
Pivotable Reverse WhoIs / PDNS Fusion、Registrant Tracking & Alerting を備え、さらに自動クエリ用API (JSON/CSV/TXT)
注意: PyDat 5 の開発中、内部の運用方針が変更され、PyDat プロジェクトは廃止されました。PyDat 5 の機能を完成させるための多くの作業は行われましたが、一部の機能は完全にテストされていません。
WhoDat プロジェクトは、whoisxmlapi データ、または ElasticSearch に保存されている任意の whois データのためのインターフェースです。whois データ、現在の IP 解決、パッシブ DNS を統合します。アナリストが調査を行うためのインタラクティブでピボット可能なアプリケーションを提供するだけでなく、JSON 形式での出力を可能にする API も備えています。
WhoDat は元々 Chris Clark によって書かれました。元の実装は PHP で、このリポジトリの legacy_whodat ディレクトリにあります。コードは Wesley Shields と Murad Khan によって Python で一から書き直され、pydat ディレクトリにあります。
PHP 版は実行したい人のために残されていますが、Python 実装ほど機能が充実しておらず、拡張性も低く、サポートされていません。
PHP 実装の詳細については、readme を参照してください。Python 実装の詳細については、以下をお読みください…
pyDat は Chris Clark の WhoDat コードの Python 実装です。PHP 実装よりも拡張性が高く、より多くの機能を備えるように設計されています。
pyDat は Python 3.6+ のアプリケーションで、実行には以下が必要です:
データベースを適切に投入するために、pydat-populator というプログラムが自動投入用に提供されています。
whoisxmlapi からのデータは常に一貫しているとは限らないため、データ取り込み時には注意が必要です。
すべてのデータが正しく取り込まれるように、さらなるテストが必要です。
データベースをセットアップする際は、スクリプトを実行する前に利用可能なフラグを読んで、自分の環境に合わせて調整してください。
以下は pydat-populator -h の出力です:
usage: pydat-populator [-h] [-c CONFIG] [--debug] [--debug-level DEBUG_LEVEL]
[-x EXCLUDE [EXCLUDE ...]] [-n INCLUDE [INCLUDE ...]]
[--ignore-field-prefixes [IGNORE_FIELD_PREFIXES [IGNORE_FIELD_PREFIXES ...]]]
[-e EXTENSION] [-v] [-s] [--pipelines PIPELINES]
[--shipper-threads SHIPPER_THREADS]
[--fetcher-threads FETCHER_THREADS]
[--bulk-ship-size BULK_SHIP_SIZE]
[--bulk-fetch-size BULK_FETCH_SIZE]
[-u [ES_URI [ES_URI ...]]] [--es-user ES_USER]
[--es-pass ES_PASSWORD] [--cacert ES_CA_CERT]
[--es-disable-sniffing] [-p ES_INDEX_PREFIX]
[--rollover-size ES_ROLLOVER_DOCS] [--ask-pass]
[-r | --config-template-only | --clear-interrupted-flag]
[-f INGEST_FILE | -d INGEST_DIRECTORY] [-D INGEST_DAY]
[-o COMMENT]
optional arguments:
-h, --help show this help message and exit
-c CONFIG, --config CONFIG
location of configuration file for
environmentparameter configuration (example yaml file
in /backend)
--debug Enables debug logging
--debug-level DEBUG_LEVEL
Debug logging level [0-3] (default: 1)
-x EXCLUDE [EXCLUDE ...], --exclude EXCLUDE [EXCLUDE ...]
list of keys to exclude if updating entry
-n INCLUDE [INCLUDE ...], --include INCLUDE [INCLUDE ...]
list of keys to include if updating entry (mutually
exclusive to -x)
--ignore-field-prefixes [IGNORE_FIELD_PREFIXES [IGNORE_FIELD_PREFIXES ...]]
list of fields (in whois data) to ignore when
extracting and inserting into ElasticSearch
-e EXTENSION, --extension EXTENSION
When scanning for CSV files only parse files with
given extension (default: csv)
-v, --verbose Be verbose
-s, --stats Print out Stats after running
-r, --redo Attempt to re-import a failed import or import more
data, uses stored metadata from previous run
--config-template-only
Configure the ElasticSearch template and then exit
--clear-interrupted-flag
Clear the interrupted flag, forcefully (NOT
RECOMMENDED)
-f INGEST_FILE, --file INGEST_FILE
Input CSV file
-d INGEST_DIRECTORY, --directory INGEST_DIRECTORY
Directory to recursively search for CSV files --
mutually exclusive to '-f' option
-D INGEST_DAY, --ingest-day INGEST_DAY
Day to use for metadata, in the format 'YYYY-MM-dd',
e.g., '2021-01-01'. Defaults to todays date, use
'YYYY-MM-00' to indicate a quarterly ingest, e.g.,
2021-04-00
-o COMMENT, --comment COMMENT
Comment to store with metadata
Performance Options:
--pipelines PIPELINES
Number of pipelines (default: 2)
--shipper-threads SHIPPER_THREADS
How many threads per pipeline to spawn to send bulk ES
messages. The larger your cluster, the more you can
increase this, defaults to 1
--fetcher-threads FETCHER_THREADS
How many threads to spawn to search ES. The larger
your cluster, the more you can increase this, defaults
to 2
--bulk-ship-size BULK_SHIP_SIZE
Size of Bulk Elasticsearch Requests (default: 10)
--bulk-fetch-size BULK_FETCH_SIZE
Number of documents to search for at a time (default:
50), note that this will be multiplied by the number
of indices you have, e.g., if you have 10
pydat-<number> indices it results in a request for 500
documents
Elasticsearch Options:
-u [ES_URI [ES_URI ...]], --es-uri [ES_URI [ES_URI ...]]
Location(s) of ElasticSearch Server (e.g.,
foo.server.com:9200) Can take multiple endpoints
--es-user ES_USER Username for ElasticSearch when Basic Auth is enabled
--es-pass ES_PASSWORD
Password for ElasticSearch when Basic Auth is enabled
--cacert ES_CA_CERT Path to a CA Certicate bundle to enable https support
--es-disable-sniffing
Disable ES sniffing, useful when ssl
hostnameverification is not working properly
-p ES_INDEX_PREFIX, --index-prefix ES_INDEX_PREFIX
Index prefix to use in ElasticSearch (default: pydat)
--rollover-size ES_ROLLOVER_DOCS
Set the number of documents after which point a new
index should be created, defaults to 50 million, note
that this is fuzzy since the index count isn't
continuously updated, so should be reasonably below 2
billion per ES shard and should take your ES
configuration into consideration
--ask-pass Prompt for ElasticSearch password
データベースに新しいバージョンのデータを追加する場合、変更を追跡する必要のない特定のフィールドを除外するために -x フラグを使用するか、監視対象の特定のフィールドのみを含めるために -n フラグを使用する必要があります。 これにより、バージョン間で保存されるデータ量が大幅に減少します。 -x または -n のいずれかのみ使用でき、両方を同時に使用することはできませんが、環境に最適な方を選択できます。 例えば、毎日更新を受け取る場合、連絡先メールアドレスが変更されたかどうかのみを気にすることに決めるかもしれませんが、四半期ごとには、重要でない特定のフィールドのみを除外したいかもしれません。
繰り返しフラグを使用する手間を省くため、pydat-populator は設定ファイルを受け付けます。
設定ファイルの作成方法の例については、サンプル設定 を参照してください。
pyDat はそれ自体ではデータを提供しません。ElasticSearch データストアに自身の whois データを提供する必要があります。