
Pivotable Reverse WhoIs / PDNS Fusion con Registrant Tracking e Alerting più API per query automatizzate (JSON/CSV/TXT)
NOTA: Durante lo sviluppo di PyDat 5, le operazioni interne hanno cambiato direzione, portando al ritiro del progetto PyDat. Sebbene gran parte del lavoro sia stato fatto per finalizzare le capacità di PyDat 5, alcune capacità rimangono non completamente testate.
Il progetto WhoDat è un'interfaccia per i dati whoisxmlapi, o per qualsiasi dato whois presente in ElasticSearch. Integra dati whois, risoluzioni IP correnti e DNS passivo. Oltre a fornire un'applicazione interattiva e orientabile per gli analisti per effettuare ricerche, dispone anche di un'API che consente l'output in formato JSON.
WhoDat è stato originariamente scritto da Chris Clark. L'implementazione originale è in PHP ed è disponibile in questo repository nella directory legacy_whodat. Il codice è stato riscritto da zero da Wesley Shields e Murad Khan in Python, ed è disponibile nella directory pydat.
La versione PHP è lasciata per coloro che vogliono eseguirla, ma non è completa o estensibile come l'implementazione Python, e non è supportata.
Per maggiori informazioni sull'implementazione PHP, consulta il readme. Per maggiori informazioni sull'implementazione Python, continua a leggere...
pyDat è un'implementazione Python del codice WhoDat di Chris Clark. È progettato per essere più estensibile e ha più funzionalità rispetto all'implementazione PHP.
pyDat è un'applicazione Python 3.6+ che richiede quanto segue per essere eseguita:
Per aiutare a popolare correttamente il database, viene fornito un programma chiamato pydat-populator per popolare automaticamente i dati. Nota che i dati provenienti da whoisxmlapi non sembrano essere sempre coerenti, quindi è necessario prestare attenzione durante l'importazione dei dati. Sono necessari ulteriori test per garantire che tutti i dati vengano importati correttamente. Chiunque configuri il proprio database dovrebbe leggere i flag disponibili per lo script prima di eseguirlo per assicurarsi di averlo adattato alla propria configurazione. Di seguito è riportato l'output di pydat-populator -h:
usage: pydat-populator [-h] [-c CONFIG] [--debug] [--debug-level DEBUG_LEVEL]
[-x EXCLUDE [EXCLUDE ...]] [-n INCLUDE [INCLUDE ...]]
[--ignore-field-prefixes [IGNORE_FIELD_PREFIXES [IGNORE_FIELD_PREFIXES ...]]]
[-e EXTENSION] [-v] [-s] [--pipelines PIPELINES]
[--shipper-threads SHIPPER_THREADS]
[--fetcher-threads FETCHER_THREADS]
[--bulk-ship-size BULK_SHIP_SIZE]
[--bulk-fetch-size BULK_FETCH_SIZE]
[-u [ES_URI [ES_URI ...]]] [--es-user ES_USER]
[--es-pass ES_PASSWORD] [--cacert ES_CA_CERT]
[--es-disable-sniffing] [-p ES_INDEX_PREFIX]
[--rollover-size ES_ROLLOVER_DOCS] [--ask-pass]
[-r | --config-template-only | --clear-interrupted-flag]
[-f INGEST_FILE | -d INGEST_DIRECTORY] [-D INGEST_DAY]
[-o COMMENT]
optional arguments:
-h, --help show this help message and exit
-c CONFIG, --config CONFIG
location of configuration file for
environmentparameter configuration (example yaml file
in /backend)
--debug Enables debug logging
--debug-level DEBUG_LEVEL
Debug logging level [0-3] (default: 1)
-x EXCLUDE [EXCLUDE ...], --exclude EXCLUDE [EXCLUDE ...]
list of keys to exclude if updating entry
-n INCLUDE [INCLUDE ...], --include INCLUDE [INCLUDE ...]
list of keys to include if updating entry (mutually
exclusive to -x)
--ignore-field-prefixes [IGNORE_FIELD_PREFIXES [IGNORE_FIELD_PREFIXES ...]]
list of fields (in whois data) to ignore when
extracting and inserting into ElasticSearch
-e EXTENSION, --extension EXTENSION
When scanning for CSV files only parse files with
given extension (default: csv)
-v, --verbose Be verbose
-s, --stats Print out Stats after running
-r, --redo Attempt to re-import a failed import or import more
data, uses stored metadata from previous run
--config-template-only
Configure the ElasticSearch template and then exit
--clear-interrupted-flag
Clear the interrupted flag, forcefully (NOT
RECOMMENDED)
-f INGEST_FILE, --file INGEST_FILE
Input CSV file
-d INGEST_DIRECTORY, --directory INGEST_DIRECTORY
Directory to recursively search for CSV files --
mutually exclusive to '-f' option
-D INGEST_DAY, --ingest-day INGEST_DAY
Day to use for metadata, in the format 'YYYY-MM-dd',
e.g., '2021-01-01'. Defaults to todays date, use
'YYYY-MM-00' to indicate a quarterly ingest, e.g.,
2021-04-00
-o COMMENT, --comment COMMENT
Comment to store with metadata
Performance Options:
--pipelines PIPELINES
Number of pipelines (default: 2)
--shipper-threads SHIPPER_THREADS
How many threads per pipeline to spawn to send bulk ES
messages. The larger your cluster, the more you can
increase this, defaults to 1
--fetcher-threads FETCHER_THREADS
How many threads to spawn to search ES. The larger
your cluster, the more you can increase this, defaults
to 2
--bulk-ship-size BULK_SHIP_SIZE
Size of Bulk Elasticsearch Requests (default: 10)
--bulk-fetch-size BULK_FETCH_SIZE
Number of documents to search for at a time (default:
50), note that this will be multiplied by the number
of indices you have, e.g., if you have 10
pydat-<number> indices it results in a request for 500
documents