
피벗 가능한 Reverse WhoIs / PDNS Fusion, 등록자 추적 및 알림, 추가로 자동 쿼리용 API (JSON/CSV/TXT)
참고: PyDat 5 개발 중 내부 작업 방향이 변경되어 PyDat 프로젝트가 중단되었습니다. PyDat 5의 기능을 완성하기 위한 많은 작업이 이루어졌지만, 일부 기능은 완전히 테스트되지 않은 상태로 남아 있습니다.
WhoDat 프로젝트는 whoisxmlapi 데이터 또는 ElasticSearch에 저장된 모든 whois 데이터를 위한 인터페이스입니다. whois 데이터, 현재 IP 확인 정보 및 수동적 DNS를 통합합니다. 분석가가 연구를 수행할 수 있는 대화형 피벗 가능한 애플리케이션을 제공할 뿐만 아니라, JSON 형식으로 출력을 제공하는 API도 갖추고 있습니다.
WhoDat은 원래 Chris Clark이 작성했습니다. 원래 구현체는 PHP로 작성되었으며, 이 리포지토리의 legacy_whodat 디렉토리에서 확인할 수 있습니다. 코드는 Wesley Shields와 Murad Khan에 의해 Python으로 처음부터 다시 작성되었으며, pydat 디렉토리에서 제공됩니다.
PHP 버전은 실행을 원하는 사람들을 위해 남겨두었지만, Python 구현체만큼 기능이 풍부하거나 확장성이 좋지 않으며 지원되지 않습니다.
PHP 구현에 대한 자세한 내용은 readme를 참조하십시오. Python 구현에 대한 자세한 내용은 계속 읽어보십시오...
pyDat는 Chris Clark의 WhoDat 코드를 Python으로 구현한 것입니다. PHP 구현보다 확장성이 뛰어나고 더 많은 기능을 제공하도록 설계되었습니다.
pyDat는 Python 3.6+ 애플리케이션으로, 실행을 위해 다음이 필요합니다:
데이터베이스를 올바르게 채우기 위해 pydat-populator라는 프로그램이 제공되어 데이터를 자동으로 채웁니다.
whoisxmlapi에서 오는 데이터가 항상 일관된 것 같지는 않으므로 데이터를 수집할 때 주의해야 합니다.
모든 데이터가 올바르게 수집되도록 더 많은 테스트가 필요합니다.
데이터베이스를 설정하는 사람은 스크립트를 실행하기 전에 사용 가능한 플래그를 읽고 자신의 설정에 맞게 조정했는지 확인해야 합니다.
다음은 pydat-populator -h의 출력입니다:
usage: pydat-populator [-h] [-c CONFIG] [--debug] [--debug-level DEBUG_LEVEL]
[-x EXCLUDE [EXCLUDE ...]] [-n INCLUDE [INCLUDE ...]]
[--ignore-field-prefixes [IGNORE_FIELD_PREFIXES [IGNORE_FIELD_PREFIXES ...]]]
[-e EXTENSION] [-v] [-s] [--pipelines PIPELINES]
[--shipper-threads SHIPPER_THREADS]
[--fetcher-threads FETCHER_THREADS]
[--bulk-ship-size BULK_SHIP_SIZE]
[--bulk-fetch-size BULK_FETCH_SIZE]
[-u [ES_URI [ES_URI ...]]] [--es-user ES_USER]
[--es-pass ES_PASSWORD] [--cacert ES_CA_CERT]
[--es-disable-sniffing] [-p ES_INDEX_PREFIX]
[--rollover-size ES_ROLLOVER_DOCS] [--ask-pass]
[-r | --config-template-only | --clear-interrupted-flag]
[-f INGEST_FILE | -d INGEST_DIRECTORY] [-D INGEST_DAY]
[-o COMMENT]
optional arguments:
-h, --help show this help message and exit
-c CONFIG, --config CONFIG
location of configuration file for
environmentparameter configuration (example yaml file
in /backend)
--debug Enables debug logging
--debug-level DEBUG_LEVEL
Debug logging level [0-3] (default: 1)
-x EXCLUDE [EXCLUDE ...], --exclude EXCLUDE [EXCLUDE ...]
list of keys to exclude if updating entry
-n INCLUDE [INCLUDE ...], --include INCLUDE [INCLUDE ...]
list of keys to include if updating entry (mutually
exclusive to -x)
--ignore-field-prefixes [IGNORE_FIELD_PREFIXES [IGNORE_FIELD_PREFIXES ...]]
list of fields (in whois data) to ignore when
extracting and inserting into ElasticSearch
-e EXTENSION, --extension EXTENSION
When scanning for CSV files only parse files with
given extension (default: csv)
-v, --verbose Be verbose
-s, --stats Print out Stats after running
-r, --redo Attempt to re-import a failed import or import more
data, uses stored metadata from previous run
--config-template-only
Configure the ElasticSearch template and then exit
--clear-interrupted-flag
Clear the interrupted flag, forcefully (NOT
RECOMMENDED)
-f INGEST_FILE, --file INGEST_FILE
Input CSV file
-d INGEST_DIRECTORY, --directory INGEST_DIRECTORY
Directory to recursively search for CSV files --
mutually exclusive to '-f' option
-D INGEST_DAY, --ingest-day INGEST_DAY
Day to use for metadata, in the format 'YYYY-MM-dd',
e.g., '2021-01-01'. Defaults to todays date, use
'YYYY-MM-00' to indicate a quarterly ingest, e.g.,
2021-04-00
-o COMMENT, --comment COMMENT
Comment to store with metadata
Performance Options:
--pipelines PIPELINES
Number of pipelines (default: 2)
--shipper-threads SHIPPER_THREADS
How many threads per pipeline to spawn to send bulk ES
messages. The larger your cluster, the more you can
increase this, defaults to 1
--fetcher-threads FETCHER_THREADS
How many threads to spawn to search ES. The larger
your cluster, the more you can increase this, defaults
to 2
--bulk-ship-size BULK_SHIP_SIZE
Size of Bulk Elasticsearch Requests (default: 10)
--bulk-fetch-size BULK_FETCH_SIZE
Number of documents to search for at a time (default:
50), note that this will be multiplied by the number
of indices you have, e.g., if you have 10
pydat-<number> indices it results in a request for 500
documents
Elasticsearch Options:
-u [ES_URI [ES_URI ...]], --es-uri [ES_URI [ES_URI ...]]
Location(s) of ElasticSearch Server (e.g.,
foo.server.com:9200) Can take multiple endpoints
--es-user ES_USER Username for ElasticSearch when Basic Auth is enabled
--es-pass ES_PASSWORD
Password for ElasticSearch when Basic Auth is enabled
--cacert ES_CA_CERT Path to a CA Certicate bundle to enable https support
--es-disable-sniffing
Disable ES sniffing, useful when ssl
hostnameverification is not working properly
-p ES_INDEX_PREFIX, --index-prefix ES_INDEX_PREFIX
Index prefix to use in ElasticSearch (default: pydat)
--rollover-size ES_ROLLOVER_DOCS
Set the number of documents after which point a new
index should be created, defaults to 50 million, note
that this is fuzzy since the index count isn't
continuously updated, so should be reasonably below 2
billion per ES shard and should take your ES
configuration into consideration
--ask-pass Prompt for ElasticSearch password
데이터베이스에 데이터의 새 버전을 추가할 때는 변경 사항을 추적하는 데 중요하지 않은 특정 필드를 제외하기 위해 -x 플래그를 사용하거나, 검토 대상이 되는 특정 필드만 포함하기 위해 -n 플래그를 사용해야 합니다. 이렇게 하면 버전 간에 저장되는 데이터 양이 크게 줄어듭니다. -x 또는 -n 중 하나만 사용할 수 있으며, 동시에 사용할 수 없습니다. 자신의 환경에 가장 적합한 것을 선택하면 됩니다. 예를 들어, 매일 업데이트를 받는 경우 contactEmail이 변경되는지만 관심이 있을 수 있지만, 분기마다는 중요하지 않다고 생각하는 특정 필드만 제외하고 싶을 수도 있습니다.
반복적인 플래그 사용 시간을 절약하기 위해 pydat-populator는 구성 파일을 사용합니다. 구성 파일을 만드는 방법에 대한 예제는 예제 구성을 참조하십시오.
pyDat는 자체적으로 데이터를 제공하지 않습니다. ElasticSearch 데이터 저장소에 자체 whois 데이터를 제공해야 합니다.