Back to updates
New releaseSep 19, 2026

urx v0.11.0

Extracts URLs from OSINT Archives for Security Insights

Share
Urx Logo

Extracts URLs from OSINT Archives for Security Insights.

Urx is a command-line tool designed for collecting URLs from OSINT archives, such as the Wayback Machine and Common Crawl. Built with Rust for efficiency, it leverages asynchronous processing to rapidly query multiple data sources. This tool simplifies the process of gathering URL information for a specified domain, providing a comprehensive dataset that can be used for various purposes, including security testing and analysis.

Features

  • Fetch URLs from multiple sources in parallel (Wayback Machine, Common Crawl, OTX, Arquivo.pt)
  • Plug in any other CDX index server — national web archives, a private pywb, OutbackCDX — with --cdx-endpoint URL, no code change needed
  • Keyless by default: Wayback, Common Crawl, OTX, Arquivo.pt, and URLScan (anonymous) all work without an API key
  • BeVigil provider: URLs extracted from unpacked Android apps — endpoints no web archive ever crawled
  • API key rotation for every keyed provider (VirusTotal, URLScan, ZoomEye, GitHub, BeVigil) to mitigate rate limits
  • Authenticated testing: -H, --cookie and --user-agent apply to every request urx makes to the target (--check-status, --extract-links, --extract-js-endpoints, --expand-specs, and the robots/sitemap probes) and are deliberately never sent to an archive
  • Filter results by file extensions, substring patterns, or full regular expressions (--match-regex / --filter-regex)
  • Predefined presets, both by file family ("no-images", "only-js") and by security interest ("only-secrets", "only-backup", "only-config", "only-api")
  • Archive-side filtering: push status code, MIME type, and date range into the CDX query itself, so filtered-out captures never cross the network
  • Client-side metadata filtering (--meta-*): filter on first/last capture date, recorded MIME type and recorded status uniformly across every provider, after collection
  • Path-scoped targets: urx example.com/shop pushes the scope into the CDX query itself (url=example.com/shop*), so a subtree of a large site costs a fraction of the whole index instead of being filtered out client-side
  • Bug-bounty scope files (--scope-file): a program's own *.example.com / !admin.example.com list used verbatim, repeatable and unioned, exclusions always winning
  • URL normalization and deduplication: Sort query parameters, remove trailing slashes, merge semantically identical URLs, and collapse near-duplicates that differ only in ids, hashes, or dates (--dedup-similar)
  • Support for multiple output formats: plain text, JSON, JSON Lines, CSV, and wordlist — the path segments and parameter names the target is built from, with ids, hashes and dates left out
  • Parameter and fuzz views: --params (the whole target's parameter inventory), --params-by-endpoint (which endpoint takes what), and --fuzz-placeholder FUZZ (one templated URL per parameter signature, ready for ffuf or dalfox)
  • Archive capture metadata: first_seen, last_seen, mime, archive_status, and digest come back with every URL a CDX archive reported, at no extra network cost
  • Streaming output (--stream): URLs are written as each provider reports them, so a pipeline starts working immediately instead of waiting for the slowest archive
  • Direct file input support: Read URLs directly from WARC files, URLTeam compressed files, and text files
  • Output results to the console or a file, or stream via stdin for pipeline integration
  • URL Testing:
    • Filter and validate URLs based on HTTP status codes and patterns.
    • Extract additional links from collected URLs — anchors, scripts, stylesheets, form actions, iframes, images, media sources, objects, embeds, and meta-refresh targets
    • Mine the archived response bodies of collected URLs (--archive-body), so pages that no longer exist still give up the links they contained — one request per distinct body, thanks to CDX digest deduplication
    • With --extract-js-endpoints, mine the archived JavaScript too: a bundle named by build hash 404s the moment the site redeploys, and the archive is the only place its API surface still exists
    • Keep the replayed bodies (--archive-body-dir) as a corpus to grep for what no link extractor looks for — developer comments, inlined credentials, internal hostnames — at no extra requests
    • Expand API specifications (--expand-specs): OpenAPI 3.x, Swagger 2.0 and GraphQL introspection documents, JSON or YAML, turned into every route they describe — one request buys the whole documented surface
    • Response metadata: --check-status also records Location, Content-Length and Content-Type, and --check-title adds the HTML <title>
  • Archived robots.txt and sitemap.xml discovery (--archived-discovery): every distinct version the Wayback Machine holds, so a Disallow: from 2015 still names the paths the site has since stopped mentioning
  • Caching and Incremental Scanning:
    • Local SQLite or remote Redis caching to avoid re-scanning domains (Redis needs a build with --features redis-cache; packaged builds leave it out)
    • Incremental mode to discover only new URLs since last scan
    • Configurable cache TTL; each scan sweeps entries older than twice the TTL, and urx cache prune removes anything past it
    • urx cache subcommand to inspect and maintain the cache: stats, list, prune, drop <domain>, clear

Preview

Installation

From Cargo

# https://crates.io/crates/urx
cargo install urx

From Homebrew

# https://formulae.brew.sh/formula/urx
brew install urx

From Source

git clone https://github.com/hahwul/urx.git
cd urx
cargo build --release

The compiled binary will be available at target/release/urx.

From Docker

ghcr.io/hahwul/urx

docker pull ghcr.io/hahwul/urx:latest
# the image has no entrypoint, so name the binary before its arguments
docker run --rm ghcr.io/hahwul/urx:latest ./urx example.com

From the AUR

yay -S urx

From Chocolatey (Windows)

choco install urx

From GitHub Releases

Prebuilt binaries for Linux, macOS and Windows (each with a .sha256) are attached to every release.

Shell Completions

urx generates its own completion script, so it always matches the flags of the binary you actually have installed.

# zsh — any directory on your $fpath works
urx --completions zsh > ~/.zfunc/_urx
# (make sure ~/.zfunc is on the fpath, then `compinit`)

Categories