
httpgrep v4.7
Scanner HTTP(S) assíncrono que busca strings ou regex em corpos de resposta e cabeçalhos em hosts, portas, intervalos CIDR e vhosts por certificado TLS.
Descrição
Uma ferramenta Python rápida e assíncrona que varre servidores HTTP(S) e procura por strings ou padrões regex nos corpos e cabeçalhos de respostas HTTP.
Ela aceita hosts únicos, URLs, faixas CIDR, faixas de IP ou arquivos; varre múltiplas portas por alvo (porta única, listas separadas por vírgula ou faixas, detectando automaticamente TLS vs. texto simples por porta); pode extrair e varrer (v)hosts baseados em nome diretamente dos certificados TLS; transmite correspondências ao vivo para o terminal; e pode gravar resultados em arquivos de log de texto, CSV ou JSONL.
Ela é feita para varreduras grandes: um núcleo assíncrono conduz milhares de conexões concorrentes, uma verificação prévia via TCP descarta portas mortas com custo baixo, timeouts por host e globais evitam que ela trave em alvos lentos/mortos, e uma execução interrompida pode ser retomada.
Requisitos
- Python 3.11+ em um sistema POSIX (Linux, *BSD, macOS - usa
termiose o tratamento de sinais Unix do asyncio) - httpx -
pip install -r requirements.txt(oupip install httpx) - opcionais, usados automaticamente se presentes:
uvloop(loop de eventos mais rápido),aiodns(dns não bloqueante para-r),h2(HTTP/2 para-2),httpx[socks]/ socksio (proxies SOCKS)
httpgrep é um script único e autocontido - basta executar ./httpgrep.py.
Uso
$ httpgrep -H
__ __ __
/ /_ / /_/ /_____ ____ _________ ____
/ __ \/ __/ __/ __ \/ __ `/ ___/ _ \/ __ \
/ / / / /_/ /_/ /_/ / /_/ / / / __/ /_/ /
/_/ /_/\__/\__/ .___/\__, /_/ \___/ .___/
/_/ /____/ /_/
--== [ by nullsecurity.net ] ==--
usage
httpgrep -h <arg> -s <arg> [opts] | <misc>
target options
-h <hosts|file> - single host/url or host-/cidr-range or file containing
hosts or file containing URLs, e.g.: foobar.net,
192.168.0.1-192.168.0.254, 192.168.0.0/24, /tmp/hosts.txt
a comma-separated list of hosts also works, e.g.:
1.2.3.4,foo.net,10.0.0.0/24
NOTE: hosts can also contain ':<ports>' on cmdline or in
file, where <ports> is a single port, comma-list or
range, e.g.: foo.net:8080, foo.net:80,443, 10.0.0.1:1-1024
-p <ports|file> - port(s) to connect to: single port, comma-separated list,
range, or a file with one spec per line, e.g.: 80,
80,443,8080, 8000-8100, /tmp/ports.txt
(default: 80, or 443 when -t is given)
-t - force TLS/SSL on all ports. by default the scheme is
auto-detected per port (plain http, switching to TLS if
the port speaks it)
-u <URI|file> - URI or comma-separated URIs or file with URIs (one per
line) to search given strings in, e.g.: /foobar/,
/foo.html, /admin,/login, /tmp/paths.txt (default: /)
-r - show the reverse-dns (PTR) name of scanned IPv4s as a
label; the ip stays the scan target (no scope drift).
non-blocking with the aiodns package
http options
-X <method> - HTTP request method to use, any case (default: get).
use '?' to list available methods.
-a <user:pass> - http auth credentials (format: 'user:pass')
-U <UA> - set custom User-Agent (default: latest ms edge, windows)
-A - use random user-agent per request
-R <headers> - set custom headers (format: 'foo=bar;lol=lulz;...')
-C <cookies> - set cookies (format: 'foo=bar;lol=lulz;...')
-F - don't follow HTTP redirects
-L <num> - max redirects to follow (default: 10; ignored with -F)
-E - verify TLS/SSL certificates (default: no verification)
-P <proxy> - use proxy (format: '[http|https|socks4|socks5]://host:port')
(socks needs the 'httpx[socks]' / socksio package)
-f <codes> - only report responses with given HTTP status codes,
e.g.: '200', '200,301,302'
-e <codes> - exclude responses with given HTTP status codes,
e.g.: '404', '403,404,500'
-2 - try HTTP/2 (ALPN-negotiated on TLS, falls back to 1.1;
plain http stays 1.1). needs the 'h2' package
search options
-s <str|file> - a single string/regex or multiple strings/regex in a file
to find in HTTP response bodies and headers (see -w),
e.g.: 'tomcat 8', '/tmp/igot0daysforthese.txt'
-S <str|file> - invert (grep -v): drop ALL matches of a response if this
string/regex (or file) appears anywhere in its body or
headers, e.g. to filter out dynamic error / 404 pages
-w <where> - where to search: headers, body, or headers,body
(default: headers,body)
-b <bytes> - num bytes of context to show from a body match
(default: 64)
-m <size> - max body to read + search; suffix b/kb/mb, no suffix = kb,
e.g.: 512, 1mb, 262144b (default: 256kb)
-i - use case-insensitive search
-I - use case-insensitive invert (for -S)
scan options
-x <num> - max concurrent connections (async; default: 300). raise
ulimit -n accordingly for very high values
-c <seconds> - per-host read timeout in seconds, also caps body read
time. the tcp preflight is capped at 2s regardless, so
filtered/dead hosts free their slot fast (default: 3.0)
-G <seconds> - global timeout: hard-stop the whole scan after N seconds
(safety net against any hang; default: none)
-y <num> - retry a failed probe up to <num> times (default: 0).
helps with flaky hosts at scale; keep it small
-1 - once a host has a match, skip its not-yet-started probes
(best-effort; in-flight requests still finish, so under
high -x you may still see a few matches per host)
-z <size> - scan targets in random order within a memory-bounded
window of <size> ram (suffix b/kb/mb/gb), e.g.: -z 1gb.
keeps huge ranges/files from exhausting memory
-Z <num> - cap the -z window at <num> targets (default 2000000,
~267mb at ~140 bytes each). more = wider mixing on huge
ranges, at the cost of ram and start-up buffering
-W - save/resume: on ctrl+c write progress to httpgrep.session;
rerun with -W to resume from it (else start fresh)
-T <0|1> - also probe the cert (v)hosts (CN + SAN) as extra requests
on top of the direct scan. 0 = via Host header on the
same ip (in-scope); 1 = ALSO by dns name/SNI (may leave
scope). needs TLS (https url, -t, or a *443 port).
output options
-l <file> - log found matches to <file>.<fmt> per chosen -O format
(e.g. -l out -O csv,jsonl => out.csv, out.jsonl)
-O <formats> - log file format(s), comma-list of: txt, csv, jsonl
(default: txt; use '?' to list). terminal output always
stays human-readable.
-v - verbose: print each url as it gets scanned
-7 - escape non-ASCII in terminal output to \xNN, so a hostile
response body can't corrupt your terminal (logs stay raw)
misc options
-H - print help
-V - print version information
examples
# grep for 'apache' in headers and body of a single host
$ httpgrep -h foobar.net -s apache
# scan a CIDR range on port 8080, search for 'tomcat' in body only
$ httpgrep -h 192.168.0.0/24 -p 8080 -s tomcat -w body
# scan a host across multiple ports and a port range for 'jenkins'
$ httpgrep -h 192.168.0.10 -p 80,443,8080,8000-8100 -s jenkins -i
# scan host list, search string file, log matches (-> /tmp/out.txt)
$ httpgrep -h /tmp/hosts.txt -s /tmp/strings.txt -x 200 -l /tmp/out
# grep for 'admin' case-insensitively across multiple URIs via TLS
$ httpgrep -h foobar.net -t -u /admin,/login,/dashboard -s admin -i
# scan IP range, reverse DNS, only report 200 responses
$ httpgrep -h 10.0.0.1-10.0.0.254 -s 'powered by' -r -f 200
# search headers only, don't follow redirects, verbose output
$ httpgrep -h foobar.net -s 'X-Powered-By' -w headers -F -v
# grep for 'admin', but drop dynamic error pages (invert, case-insensitive)
$ httpgrep -h 192.168.0.0/24 -s admin -i -S 'error|not found' -I
# route through proxy, custom UA, search for version strings
$ httpgrep -h /tmp/hosts.txt -s 'nginx/1\.' -P http://127.0.0.1:8080 -U 'curl/8.0'
# big resumable scan: ctrl+c saves state, rerun with -W to continue; also
# cap the whole run at 1 hour as a hang safety net
$ httpgrep -h 10.0.0.0/16 -p 80,443 -s admin -W -G 3600
Saída
As correspondências são exibidas ao vivo, uma por linha:
[*] <url> | [vhost] | <status> | <type> | <match>
<url>- a URL varrida (scheme://host:port/uri).<vhost>- presente com-T(o (v)host do certificado tentado via cabeçalhoHost) ou-r(o nome PTR do IP varrido); vazio para varreduras diretas.<status>- o código de status HTTP da resposta (após redirecionamentos).<type>-bodyouheader.<match>- acerto no corpo: uma pequena janela repr do trecho correspondente (-bbytes); acerto no cabeçalho:nome: valor.
O terminal sempre mostra essa forma legível. Com -l <base>, as mesmas correspondências são espelhadas em <base>.<fmt> para cada formato -O - txt (essas linhas), csv (linhas url,vhost,status,type,match com cabeçalho), jsonl (um objeto JSON por correspondência).
Em uma varredura com múltiplos alvos, uma linha de status ao vivo é exibida (fixada na parte inferior em um tty, linhas simples quando redirecionado):
[+] wait bitch, scanning: <targets> | <scanned>/<total> | <pct>% | <n> hits
<total> é a contagem de alvos (cidrs/faixas calculados, não expandidos); <n> hits é o número acumulado de linhas de correspondência emitidas até agora.
Autor
noptrix
Notas
- código rápido e sujo
- httpgrep já está empacotado e disponível para BlackArch Linux
- Meus branches master são sempre estáveis; branches de desenvolvimento são criados para o trabalho atual.
- Tudo o que você encontrar de meu material público é oficialmente anunciado e publicado via nullsecurity.net.
Licença
Consulte docs/LICENSE.
Aviso Legal
Enfatizamos aqui que o material relacionado a hacking encontrado em nullsecurity.net é apenas para fins educacionais. Não somos responsáveis por quaisquer danos. Você é responsável pelas suas próprias ações.