アップデート一覧に戻る
New releaseAug 5, 2026

katana v1.7.0

次世代のクローリングおよびスパイダリングフレームワーク。

共有

katana

次世代のクローリングおよびスパイダリングフレームワーク

機能 • インストール • 使用方法 • スコープ • 設定 • フィルター • Discordに参加

機能

image

  • 高速かつ完全に設定可能なWebクローリング
  • StandardモードとHeadlessモード
  • JavaScriptの解析 / クローリング
  • カスタマイズ可能な自動フォーム入力
  • スコープ制御 - 事前設定されたフィールド / 正規表現
  • ナレッジベース - MLによるページタイプ / フォーム分類(モデルは自動ダウンロード)
  • カスタマイズ可能な出力 - 事前設定されたフィールド
  • 入力 - STDIN、URL、LIST
  • 出力 - STDOUT、FILE、JSON

インストール

katanaを正常にインストールするにはGo 1.26以上が必要です。インストール時に問題が発生した場合は、最低要件のバージョンが変更されている可能性があるため、利用可能な最新バージョンのGoで試すことを推奨します。以下のコマンドを実行するか、リリースページからコンパイル済みバイナリをダウンロードしてください。```console CGO_ENABLED=1 go install github.com/projectdiscovery/katana/cmd/katana@latest

**katana をインストール / 実行するためのその他のオプション**

<details>
  <summary>Docker</summary>

> docker を最新タグにインストール / 更新するには -```sh
docker pull projectdiscovery/katana:latest

Docker を使用して katana を標準モードで実行するには -```sh docker run projectdiscovery/katana:latest -u https://tesla.com

> Docker を使用して katana をヘッドレスモードで実行するには -```sh
docker run projectdiscovery/katana:latest -u https://tesla.com -system-chrome -headless
Ubuntu

以下の前提条件をインストールすることを推奨します -```sh sudo apt update sudo apt install zip curl wget git snapd sudo snap refresh sudo snap install golang --classic

sudo install -d -m 0755 /etc/apt/keyrings curl -fsSL https://dl.google.com/linux/linux_signing_key.pub
| sudo gpg --dearmor -o /etc/apt/keyrings/google-chrome.gpg

echo "deb [arch=amd64 signed-by=/etc/apt/keyrings/google-chrome.gpg]
http://dl.google.com/linux/chrome/deb/ stable main"
| sudo tee /etc/apt/sources.list.d/google-chrome.list > /dev/null

sudo apt update sudo apt install google-chrome-stable

> katana をインストール -```sh
go install github.com/projectdiscovery/katana/cmd/katana@latest

使用方法```console

katana -h

これはツールのヘルプを表示します。サポートされているすべてのスイッチは以下のとおりです。```console
Katana is a fast crawler focused on execution in automation
pipelines offering both headless and non-headless crawling.

Usage:
  ./katana [flags]

Flags:
INPUT:
   -u, -list string[]     target url / list to crawl
   -resume string         resume scan using resume.cfg
   -e, -exclude string[]  exclude host matching specified filter ('cdn', 'private-ips', cidr, ip, regex)

CONFIGURATION:
   -r, -resolvers string[]       list of custom resolver (file or comma separated)
   -d, -depth int                maximum depth to crawl (default 3)
   -jc, -js-crawl                enable endpoint parsing / crawling in javascript file
   -jsl, -jsluice                enable jsluice parsing in javascript file (memory intensive)
   -ct, -crawl-duration value    maximum duration to crawl the target for (s, m, h, d) (default s)
   -kf, -known-files string      enable crawling of known files (all,robotstxt,sitemapxml), a minimum depth of 3 is required to ensure all known files are properly crawled.
   -mrs, -max-response-size int  maximum response size to read (default 4194304)
   -timeout int                  time to wait for request in seconds (default 10)
   -aff, -automatic-form-fill    enable automatic form filling (experimental)
   -fx, -form-extraction         extract form, input, textarea & select elements in jsonl output
   -retry int                    number of times to retry the request (default 1)
   -proxy string                 http/socks5 proxy to use
   -td, -tech-detect             enable technology detection
   -H, -headers string[]         custom header/cookie to include in all http request in header:value format (file)
   -config string                path to the katana configuration file
   -fc, -form-config string      path to custom form configuration file
   -flc, -field-config string    path to custom field configuration file
   -s, -strategy string          Visit strategy (depth-first, breadth-first) (default "depth-first")
   -iqp, -ignore-query-params    Ignore crawling same path with different query-param values
   -fsu, -filter-similar         filter crawling of similar looking URLs (e.g., /users/123 and /users/456)
   -fst, -filter-similar-threshold int  number of distinct values before a path position is treated as parameter (default 10)
   -tlsi, -tls-impersonate       enable experimental client hello (ja3) tls randomization
   -dr, -disable-redirects       disable following redirects (default false)
   -pcs, -page-content-similar   enable page content similarity filtering (simhash|tfidf|bm25)
   -pcsm, -page-content-similar-mode string  similarity mode: simhash, tfidf, or bm25 (default simhash)
   -pcsd, -page-content-similar-distance int  simhash max hamming distance (default 3)
   -pcst, -page-content-similar-threshold float  tfidf/bm25 min score 0-1 (default 0.85)
   -pcsn, -page-content-similar-budget int  pages to fully process per similarity cluster (default 1)
   -sdd, -similarity-deduplication  alias for -pcs
   -kb, -knowledge-base          enable knowledge base classification
   -kb-secrets                   enable secrets extractor in the knowledge base
   -kb-validate-secrets          validate detected secrets against their provider (sends live API calls)
   -kb-endpoints                 enable endpoints extractor (classifies REST/GraphQL/SOAP/XHR requests)
   -mdp, -max-domain-pages int   maximum number of pages to crawl per domain (default unlimited)

DEBUG:
   -health-check, -hc        run diagnostic check up
   -elog, -error-log string  file to write sent requests error log
   -pprof-server             enable pprof server

カテゴリ