
WhoIs Reverso Pivotável / Fusão PDNS com Rastreamento do Titular e Alertas, além de API para consultas automatizadas (JSON/CSV/TXT)
NOTA: Durante o desenvolvimento do PyDat 5, as operações internas mudaram de direção, levando à aposentadoria do projeto PyDat. Embora muito trabalho tenha sido feito para finalizar as capacidades do PyDat 5, algumas capacidades ainda não foram totalmente testadas.
O projeto WhoDat é uma interface para dados whoisxmlapi, ou qualquer dado whois armazenado no ElasticSearch. Ele integra dados whois, resoluções IP atuais e DNS passivo. Além de fornecer um aplicativo interativo e pivotável para analistas realizarem pesquisas, também possui uma API que permite saída em formato JSON.
WhoDat foi originalmente escrito por Chris Clark. A implementação original está em PHP e está disponível neste repositório no diretório legacy_whodat. O código foi reescrito do zero por Wesley Shields e Murad Khan em Python, e está disponível no diretório pydat.
A versão PHP é deixada para quem quiser executá-la, mas não é tão completa ou extensível quanto a implementação Python, e não é suportada.
Para mais informações sobre a implementação PHP, consulte o readme. Para mais informações sobre a implementação Python, continue lendo...
pyDat é uma implementação em Python do código WhoDat de Chris Clark. Ela foi projetada para ser mais extensível e possui mais recursos do que a implementação PHP.
pyDat é um aplicativo Python 3.6+ que requer o seguinte para ser executado:
Para ajudar a preencher corretamente o banco de dados, um programa chamado pydat-populator é fornecido para auto-povoar os dados.
Observe que os dados provenientes de whoisxmlapi parecem nem sempre ser consistentes, portanto, deve-se ter cuidado ao ingerir dados.
Mais testes precisam ser feitos para garantir que todos os dados sejam ingeridos corretamente.
Qualquer pessoa que configure seu banco de dados deve ler as flags disponíveis para o script antes de executá-lo para garantir que o ajustou para sua configuração.
A seguir está a saída de pydat-populator -h:
usage: pydat-populator [-h] [-c CONFIG] [--debug] [--debug-level DEBUG_LEVEL]
[-x EXCLUDE [EXCLUDE ...]] [-n INCLUDE [INCLUDE ...]]
[--ignore-field-prefixes [IGNORE_FIELD_PREFIXES [IGNORE_FIELD_PREFIXES ...]]]
[-e EXTENSION] [-v] [-s] [--pipelines PIPELINES]
[--shipper-threads SHIPPER_THREADS]
[--fetcher-threads FETCHER_THREADS]
[--bulk-ship-size BULK_SHIP_SIZE]
[--bulk-fetch-size BULK_FETCH_SIZE]
[-u [ES_URI [ES_URI ...]]] [--es-user ES_USER]
[--es-pass ES_PASSWORD] [--cacert ES_CA_CERT]
[--es-disable-sniffing] [-p ES_INDEX_PREFIX]
[--rollover-size ES_ROLLOVER_DOCS] [--ask-pass]
[-r | --config-template-only | --clear-interrupted-flag]
[-f INGEST_FILE | -d INGEST_DIRECTORY] [-D INGEST_DAY]
[-o COMMENT]
optional arguments:
-h, --help show this help message and exit
-c CONFIG, --config CONFIG
location of configuration file for
environmentparameter configuration (example yaml file
in /backend)
--debug Enables debug logging
--debug-level DEBUG_LEVEL
Debug logging level [0-3] (default: 1)
-x EXCLUDE [EXCLUDE ...], --exclude EXCLUDE [EXCLUDE ...]
list of keys to exclude if updating entry
-n INCLUDE [INCLUDE ...], --include INCLUDE [INCLUDE ...]
list of keys to include if updating entry (mutually
exclusive to -x)
--ignore-field-prefixes [IGNORE_FIELD_PREFIXES [IGNORE_FIELD_PREFIXES ...]]
list of fields (in whois data) to ignore when
extracting and inserting into ElasticSearch
-e EXTENSION, --extension EXTENSION
When scanning for CSV files only parse files with
given extension (default: csv)
-v, --verbose Be verbose
-s, --stats Print out Stats after running
-r, --redo Attempt to re-import a failed import or import more
data, uses stored metadata from previous run
--config-template-only
Configure the ElasticSearch template and then exit
--clear-interrupted-flag
Clear the interrupted flag, forcefully (NOT
RECOMMENDED)
-f INGEST_FILE, --file INGEST_FILE
Input CSV file
-d INGEST_DIRECTORY, --directory INGEST_DIRECTORY
Directory to recursively search for CSV files --
mutually exclusive to '-f' option
-D INGEST_DAY, --ingest-day INGEST_DAY
Day to use for metadata, in the format 'YYYY-MM-dd',
e.g., '2021-01-01'. Defaults to todays date, use
'YYYY-MM-00' to indicate a quarterly ingest, e.g.,
2021-04-00
-o COMMENT, --comment COMMENT
Comment to store with metadata
Performance Options:
--pipelines PIPELINES
Number of pipelines (default: 2)
--shipper-threads SHIPPER_THREADS
How many threads per pipeline to spawn to send bulk ES
messages. The larger your cluster, the more you can
increase this, defaults to 1
--fetcher-threads FETCHER_THREADS
How many threads to spawn to search ES. The larger
your cluster, the more you can increase this, defaults
to 2
--bulk-ship-size BULK_SHIP_SIZE
Size of Bulk Elasticsearch Requests (default: 10)
--bulk-fetch-size BULK_FETCH_SIZE
Number of documents to search for at a time (default:
50), note that this will be multiplied by the number
of indices you have, e.g., if you have 10
pydat-<number> indices it results in a request for 500
documents