
katana v1.7.0
Uma estrutura de crawling e spidering de próxima geração.
Um framework de crawling e spidering de próxima geração
Recursos • Instalação • Uso • Escopo • Configuração • Filtros • Entre no Discord
Recursos

- Web crawling rápido e totalmente configurável
- Modo Standard e Headless
- Análise / crawling de JavaScript
- Preenchimento automático de formulários personalizável
- Controle de escopo - Campo / Regex pré-configurado
- Base de conhecimento - Classificação de tipo de página / formulário por ML (modelo baixado automaticamente)
- Saída personalizável - Campos pré-configurados
- ENTRADA - STDIN, URL e LIST
- SAÍDA - STDOUT, FILE e JSON
Instalação
katana requer Go 1.26+ para ser instalado com sucesso. Se você encontrar algum problema de instalação, recomendamos tentar com a versão mais recente disponível do Go, pois a versão mínima necessária pode ter mudado. Execute o comando abaixo ou baixe um binário pré-compilado da página de releases.```console CGO_ENABLED=1 go install github.com/projectdiscovery/katana/cmd/katana@latest
**Mais opções para instalar / executar o katana-**
<details>
<summary>Docker</summary>
> Para instalar / atualizar o docker para a tag mais recente -```sh
docker pull projectdiscovery/katana:latest
Para executar o katana no modo padrão usando docker -```sh docker run projectdiscovery/katana:latest -u https://tesla.com
> Para executar o katana em modo headless usando docker -```sh
docker run projectdiscovery/katana:latest -u https://tesla.com -system-chrome -headless
Ubuntu
Recomenda-se instalar os seguintes pré-requisitos -```sh sudo apt update sudo apt install zip curl wget git snapd sudo snap refresh sudo snap install golang --classic
sudo install -d -m 0755 /etc/apt/keyrings
curl -fsSL https://dl.google.com/linux/linux_signing_key.pub
| sudo gpg --dearmor -o /etc/apt/keyrings/google-chrome.gpg
echo "deb [arch=amd64 signed-by=/etc/apt/keyrings/google-chrome.gpg]
http://dl.google.com/linux/chrome/deb/ stable main"
| sudo tee /etc/apt/sources.list.d/google-chrome.list > /dev/null
sudo apt update sudo apt install google-chrome-stable
> instalar katana -```sh
go install github.com/projectdiscovery/katana/cmd/katana@latest
Uso```console
katana -h
Isso exibirá a ajuda para a ferramenta. Aqui estão todas as opções que ela suporta.```console
Katana is a fast crawler focused on execution in automation
pipelines offering both headless and non-headless crawling.
Usage:
./katana [flags]
Flags:
INPUT:
-u, -list string[] target url / list to crawl
-resume string resume scan using resume.cfg
-e, -exclude string[] exclude host matching specified filter ('cdn', 'private-ips', cidr, ip, regex)
CONFIGURATION:
-r, -resolvers string[] list of custom resolver (file or comma separated)
-d, -depth int maximum depth to crawl (default 3)
-jc, -js-crawl enable endpoint parsing / crawling in javascript file
-jsl, -jsluice enable jsluice parsing in javascript file (memory intensive)
-ct, -crawl-duration value maximum duration to crawl the target for (s, m, h, d) (default s)
-kf, -known-files string enable crawling of known files (all,robotstxt,sitemapxml), a minimum depth of 3 is required to ensure all known files are properly crawled.
-mrs, -max-response-size int maximum response size to read (default 4194304)
-timeout int time to wait for request in seconds (default 10)
-aff, -automatic-form-fill enable automatic form filling (experimental)
-fx, -form-extraction extract form, input, textarea & select elements in jsonl output
-retry int number of times to retry the request (default 1)
-proxy string http/socks5 proxy to use
-td, -tech-detect enable technology detection
-H, -headers string[] custom header/cookie to include in all http request in header:value format (file)
-config string path to the katana configuration file
-fc, -form-config string path to custom form configuration file
-flc, -field-config string path to custom field configuration file
-s, -strategy string Visit strategy (depth-first, breadth-first) (default "depth-first")
-iqp, -ignore-query-params Ignore crawling same path with different query-param values
-fsu, -filter-similar filter crawling of similar looking URLs (e.g., /users/123 and /users/456)
-fst, -filter-similar-threshold int number of distinct values before a path position is treated as parameter (default 10)
-tlsi, -tls-impersonate enable experimental client hello (ja3) tls randomization
-dr, -disable-redirects disable following redirects (default false)
-pcs, -page-content-similar enable page content similarity filtering (simhash|tfidf|bm25)
-pcsm, -page-content-similar-mode string similarity mode: simhash, tfidf, or bm25 (default simhash)
-pcsd, -page-content-similar-distance int simhash max hamming distance (default 3)
-pcst, -page-content-similar-threshold float tfidf/bm25 min score 0-1 (default 0.85)
-pcsn, -page-content-similar-budget int pages to fully process per similarity cluster (default 1)
-sdd, -similarity-deduplication alias for -pcs
-kb, -knowledge-base enable knowledge base classification
-kb-secrets enable secrets extractor in the knowledge base
-kb-validate-secrets validate detected secrets against their provider (sends live API calls)
-kb-endpoints enable endpoints extractor (classifies REST/GraphQL/SOAP/XHR requests)
-mdp, -max-domain-pages int maximum number of pages to crawl per domain (default unlimited)
DEBUG:
-health-check, -hc run diagnostic check up
-elog, -error-log string file to write sent requests error log
-pprof-server enable pprof server