Skip to content
KitploitKITPLOIT
ToolsBlog
Log in
Submit
ToolsBlog
Submit

Hacking, PenTest, and Cybersecurity Tools for Your Security Arsenal!

Kitploit is a directory of hacking, cybersecurity, and pentesting tools. Discover the latest project updates to find vulnerabilities, analyze systems, automate testing, and strengthen your security.

··Feeds·Contact·Privacy·© 2026 Kitploit

Tool Directory

Categories

View all categories
Loading categories
urx — Extracts URLs from OSINT Archives for Security Insights | Kitploit
Tools/GitHubGitHub/hahwul/urx
OSINT (Open Source Intelligence)ReconnaissanceInformation GatheringWeb SecurityCrawler
GitHubhahwul/urx

urx

Extracts URLs from OSINT Archives for Security Insights

View Repository
19020495 days agoReviewed by Kitploit

Most Popular

View all →

Discover the most used tools by our community.

Explore all tools

Browse our collection of tools

View all tools →
Share
Website
Urx Logo

Extracts URLs from OSINT Archives for Security Insights.

Urx is a command-line tool designed for collecting URLs from OSINT archives, such as the Wayback Machine and Common Crawl. Built with Rust for efficiency, it leverages asynchronous processing to rapidly query multiple data sources. This tool simplifies the process of gathering URL information for a specified domain, providing a comprehensive dataset that can be used for various purposes, including security testing and analysis.

Features

  • Fetch URLs from multiple sources in parallel (Wayback Machine, Common Crawl, OTX, Arquivo.pt)
  • Plug in any other CDX index server — national web archives, a private pywb, OutbackCDX — with --cdx-endpoint URL, no code change needed
  • Keyless by default: Wayback, Common Crawl, OTX, Arquivo.pt, and URLScan (anonymous) all work without an API key
  • BeVigil provider: URLs extracted from unpacked Android apps — endpoints no web archive ever crawled
  • API key rotation support for VirusTotal and URLScan providers to mitigate rate limits
  • Authenticated testing: -H, --cookie and --user-agent apply to every request urx makes to the target (--check-status, --extract-links, --extract-js-endpoints, --expand-specs) and are deliberately never sent to an archive
  • Filter results by file extensions, substring patterns, or full regular expressions (--match-regex / --filter-regex)
  • Predefined presets, both by file family ("no-images", "only-js") and by security interest ("only-secrets", "only-backup", "only-config", "only-api")
  • Archive-side filtering: push status code, MIME type, and date range into the CDX query itself, so filtered-out captures never cross the network
  • Client-side metadata filtering (--meta-*): filter on first/last capture date, recorded MIME type and recorded status uniformly across every provider, after collection
  • Path-scoped targets: urx example.com/shop pushes the scope into the CDX query itself (url=example.com/shop*), so a subtree of a large site costs a fraction of the whole index instead of being filtered out client-side
  • Bug-bounty scope files (--scope-file): a program's own *.example.com / !admin.example.com list used verbatim, repeatable and unioned, exclusions always winning
  • URL normalization and deduplication: Sort query parameters, remove trailing slashes, merge semantically identical URLs, and collapse near-duplicates that differ only in ids, hashes, or dates (--dedup-similar)
  • Support for multiple output formats: plain text, JSON, JSON Lines, CSV, and wordlist — the path segments and parameter names the target is built from, with ids, hashes and dates left out
  • Parameter and fuzz views: --params (the whole target's parameter inventory), --params-by-endpoint (which endpoint takes what), and --fuzz-placeholder FUZZ (one templated URL per parameter signature, ready for ffuf or dalfox)
  • Archive capture metadata: first_seen, last_seen, mime, archive_status, and digest come back with every URL a CDX archive reported, at no extra network cost
  • Streaming output (--stream): URLs are written as each provider reports them, so a pipeline starts working immediately instead of waiting for the slowest archive
  • Direct file input support: Read URLs directly from WARC files, URLTeam compressed files, and text files
  • Output results to the console or a file, or stream via stdin for pipeline integration
  • URL Testing:
    • Filter and validate URLs based on HTTP status codes and patterns.
    • Extract additional links from collected URLs — anchors, scripts, stylesheets, form actions, iframes, images, media sources, objects, embeds, and meta-refresh targets
    • Mine the archived response bodies of collected URLs (--archive-body), so pages that no longer exist still give up the links they contained — one request per distinct body, thanks to CDX digest deduplication
    • With --extract-js-endpoints, mine the archived JavaScript too: a bundle named by build hash 404s the moment the site redeploys, and the archive is the only place its API surface still exists
    • Keep the replayed bodies (--archive-body-dir) as a corpus to grep for what no link extractor looks for — developer comments, inlined credentials, internal hostnames — at no extra requests
    • Expand API specifications (--expand-specs): OpenAPI 3.x, Swagger 2.0 and GraphQL introspection documents, JSON or YAML, turned into every route they describe — one request buys the whole documented surface
    • Response metadata: --check-status also records Location, Content-Length and Content-Type, and --check-title adds the HTML <title>
  • Archived robots.txt and sitemap.xml discovery (--archived-discovery): every distinct version the Wayback Machine holds, so a Disallow: from 2015 still names the paths the site has since stopped mentioning
  • Caching and Incremental Scanning:
    • Local SQLite or remote Redis caching to avoid re-scanning domains
    • Incremental mode to discover only new URLs since last scan
    • Configurable cache TTL and automatic cleanup of expired entries
    • urx cache subcommand to inspect and maintain the cache: stats, list, prune, drop <domain>, clear

Preview

Installation

From Cargo

# https://crates.io/crates/urx
cargo install urx

From Homebrew

# https://formulae.brew.sh/formula/urx
brew install urx

From Source

git clone https://github.com/hahwul/urx.git
cd urx
cargo build --release

The compiled binary will be available at target/release/urx.

From Docker

ghcr.io/hahwul/urx

Shell Completions

urx generates its own completion script, so it always matches the flags of the binary you actually have installed.

# zsh — any directory on your $fpath works
urx --completions zsh > ~/.zfunc/_urx
# (make sure ~/.zfunc is on the fpath, then `compinit`)

# bash
urx --completions bash > ~/.local/share/bash-completion/completions/urx

# fish
urx --completions fish > ~/.config/fish/completions/urx.fish

powershell and elvish are supported too. The flag needs no target domain.

Man Page

urx --manpage > ~/.local/share/man/man1/urx.1
man urx

Usage

Basic Usage

# Scan a single domain
urx example.com

# Scan multiple domains
urx example.com example.org

# Scan domains from a file
cat domains.txt | urx

Options

Usage: urx [OPTIONS] [DOMAINS]... [COMMAND]

Commands:
  cache  Inspect and maintain the URL cache: stats, list, prune, drop <DOMAIN>..., clear
Download Tool