
urx v0.11.0
Extracts URLs from OSINT Archives for Security Insights
Extracts URLs from OSINT Archives for Security Insights.
Urx is a command-line tool designed for collecting URLs from OSINT archives, such as the Wayback Machine and Common Crawl. Built with Rust for efficiency, it leverages asynchronous processing to rapidly query multiple data sources. This tool simplifies the process of gathering URL information for a specified domain, providing a comprehensive dataset that can be used for various purposes, including security testing and analysis.
Features
- Fetch URLs from multiple sources in parallel (Wayback Machine, Common Crawl, OTX, Arquivo.pt)
- Plug in any other CDX index server — national web archives, a private pywb, OutbackCDX — with
--cdx-endpoint URL, no code change needed - Keyless by default: Wayback, Common Crawl, OTX, Arquivo.pt, and URLScan (anonymous) all work without an API key
- BeVigil provider: URLs extracted from unpacked Android apps — endpoints no web archive ever crawled
- API key rotation for every keyed provider (VirusTotal, URLScan, ZoomEye, GitHub, BeVigil) to mitigate rate limits
- Authenticated testing:
-H,--cookieand--user-agentapply to every request urx makes to the target (--check-status,--extract-links,--extract-js-endpoints,--expand-specs, and therobots/sitemapprobes) and are deliberately never sent to an archive - Filter results by file extensions, substring patterns, or full regular expressions (
--match-regex/--filter-regex) - Predefined presets, both by file family ("no-images", "only-js") and by security interest ("only-secrets", "only-backup", "only-config", "only-api")
- Archive-side filtering: push status code, MIME type, and date range into the CDX query itself, so filtered-out captures never cross the network
- Client-side metadata filtering (
--meta-*): filter on first/last capture date, recorded MIME type and recorded status uniformly across every provider, after collection - Path-scoped targets:
urx example.com/shoppushes the scope into the CDX query itself (url=example.com/shop*), so a subtree of a large site costs a fraction of the whole index instead of being filtered out client-side - Bug-bounty scope files (
--scope-file): a program's own*.example.com/!admin.example.comlist used verbatim, repeatable and unioned, exclusions always winning - URL normalization and deduplication: Sort query parameters, remove trailing slashes, merge semantically identical URLs, and collapse near-duplicates that differ only in ids, hashes, or dates (
--dedup-similar) - Support for multiple output formats: plain text, JSON, JSON Lines, CSV, and
wordlist— the path segments and parameter names the target is built from, with ids, hashes and dates left out - Parameter and fuzz views:
--params(the whole target's parameter inventory),--params-by-endpoint(which endpoint takes what), and--fuzz-placeholder FUZZ(one templated URL per parameter signature, ready for ffuf or dalfox) - Archive capture metadata:
first_seen,last_seen,mime,archive_status, anddigestcome back with every URL a CDX archive reported, at no extra network cost - Streaming output (
--stream): URLs are written as each provider reports them, so a pipeline starts working immediately instead of waiting for the slowest archive - Direct file input support: Read URLs directly from WARC files, URLTeam compressed files, and text files
- Output results to the console or a file, or stream via stdin for pipeline integration
- URL Testing:
- Filter and validate URLs based on HTTP status codes and patterns.
- Extract additional links from collected URLs — anchors, scripts, stylesheets, form actions, iframes, images, media sources, objects, embeds, and meta-refresh targets
- Mine the archived response bodies of collected URLs (
--archive-body), so pages that no longer exist still give up the links they contained — one request per distinct body, thanks to CDX digest deduplication - With
--extract-js-endpoints, mine the archived JavaScript too: a bundle named by build hash 404s the moment the site redeploys, and the archive is the only place its API surface still exists - Keep the replayed bodies (
--archive-body-dir) as a corpus to grep for what no link extractor looks for — developer comments, inlined credentials, internal hostnames — at no extra requests - Expand API specifications (
--expand-specs): OpenAPI 3.x, Swagger 2.0 and GraphQL introspection documents, JSON or YAML, turned into every route they describe — one request buys the whole documented surface - Response metadata:
--check-statusalso recordsLocation,Content-LengthandContent-Type, and--check-titleadds the HTML<title>
- Archived robots.txt and sitemap.xml discovery (
--archived-discovery): every distinct version the Wayback Machine holds, so aDisallow:from 2015 still names the paths the site has since stopped mentioning - Caching and Incremental Scanning:
- Local SQLite or remote Redis caching to avoid re-scanning domains (Redis needs a build with
--features redis-cache; packaged builds leave it out) - Incremental mode to discover only new URLs since last scan
- Configurable cache TTL; each scan sweeps entries older than twice the TTL, and
urx cache pruneremoves anything past it urx cachesubcommand to inspect and maintain the cache:stats,list,prune,drop <domain>,clear
- Local SQLite or remote Redis caching to avoid re-scanning domains (Redis needs a build with

Installation
From Cargo
# https://crates.io/crates/urx
cargo install urx
From Homebrew
# https://formulae.brew.sh/formula/urx
brew install urx
From Source
git clone https://github.com/hahwul/urx.git
cd urx
cargo build --release
The compiled binary will be available at target/release/urx.
From Docker
docker pull ghcr.io/hahwul/urx:latest
# the image has no entrypoint, so name the binary before its arguments
docker run --rm ghcr.io/hahwul/urx:latest ./urx example.com
From the AUR
yay -S urx
From Chocolatey (Windows)
choco install urx
From GitHub Releases
Prebuilt binaries for Linux, macOS and Windows (each with a .sha256) are
attached to every release.
Shell Completions
urx generates its own completion script, so it always matches the flags of
the binary you actually have installed.
# zsh — any directory on your $fpath works
urx --completions zsh > ~/.zfunc/_urx
# (make sure ~/.zfunc is on the fpath, then `compinit`)