
Pivotable Reverse WhoIs / Fusión PDNS con Seguimiento de Titulares y Alertas más API para consultas automatizadas (JSON/CSV/TXT)
NOTA: Durante el desarrollo de PyDat 5, las operaciones internas cambiaron su dirección, lo que llevó al retiro del proyecto PyDat. Aunque se ha realizado gran parte del trabajo para finalizar las capacidades de PyDat 5, algunas capacidades aún no se han probado completamente.
El proyecto WhoDat es una interfaz para datos de whoisxmlapi, o cualquier dato whois que resida en ElasticSearch. Integra datos whois, resoluciones de IP actuales y DNS pasivo. Además de proporcionar una aplicación interactiva y pivotable para que los analistas realicen investigaciones, también tiene una API que permitirá la salida en formato JSON.
WhoDat fue originalmente escrito por Chris Clark. La implementación original está en PHP y está disponible en este repositorio bajo el directorio legacy_whodat. El código fue reescrito desde cero por Wesley Shields y Murad Khan en Python, y está disponible bajo el directorio pydat.
La versión PHP se deja para aquellos que quieran ejecutarla, pero no tiene tantas funciones ni es tan extensible como la implementación en Python, y no tiene soporte.
Para más información sobre la implementación en PHP, consulte el readme. Para más información sobre la implementación en Python, siga leyendo...
pyDat es una implementación en Python del código WhoDat de Chris Clark. Está diseñado para ser más extensible y tener más características que la implementación en PHP.
pyDat es una aplicación de Python 3.6+ que requiere lo siguiente para ejecutarse:
Para ayudar a poblar correctamente la base de datos, se proporciona un programa llamado pydat-populator para auto-poblar los datos. Tenga en cuenta que los datos provenientes de whoisxmlapi no parecen ser siempre consistentes, por lo que se debe tener cuidado al ingerir datos. Se necesita realizar más pruebas para garantizar que todos los datos se ingieran correctamente. Cualquier persona que configure su base de datos debe leer las banderas disponibles para el script antes de ejecutarlo para asegurarse de haberlo ajustado a su configuración. A continuación se muestra la salida de pydat-populator -h:
usage: pydat-populator [-h] [-c CONFIG] [--debug] [--debug-level DEBUG_LEVEL]
[-x EXCLUDE [EXCLUDE ...]] [-n INCLUDE [INCLUDE ...]]
[--ignore-field-prefixes [IGNORE_FIELD_PREFIXES [IGNORE_FIELD_PREFIXES ...]]]
[-e EXTENSION] [-v] [-s] [--pipelines PIPELINES]
[--shipper-threads SHIPPER_THREADS]
[--fetcher-threads FETCHER_THREADS]
[--bulk-ship-size BULK_SHIP_SIZE]
[--bulk-fetch-size BULK_FETCH_SIZE]
[-u [ES_URI [ES_URI ...]]] [--es-user ES_USER]
[--es-pass ES_PASSWORD] [--cacert ES_CA_CERT]
[--es-disable-sniffing] [-p ES_INDEX_PREFIX]
[--rollover-size ES_ROLLOVER_DOCS] [--ask-pass]
[-r | --config-template-only | --clear-interrupted-flag]
[-f INGEST_FILE | -d INGEST_DIRECTORY] [-D INGEST_DAY]
[-o COMMENT]
optional arguments:
-h, --help show this help message and exit
-c CONFIG, --config CONFIG
location of configuration file for
environmentparameter configuration (example yaml file
in /backend)
--debug Enables debug logging
--debug-level DEBUG_LEVEL
Debug logging level [0-3] (default: 1)
-x EXCLUDE [EXCLUDE ...], --exclude EXCLUDE [EXCLUDE ...]
list of keys to exclude if updating entry
-n INCLUDE [INCLUDE ...], --include INCLUDE [INCLUDE ...]
list of keys to include if updating entry (mutually
exclusive to -x)
--ignore-field-prefixes [IGNORE_FIELD_PREFIXES [IGNORE_FIELD_PREFIXES ...]]
list of fields (in whois data) to ignore when
extracting and inserting into ElasticSearch
-e EXTENSION, --extension EXTENSION
When scanning for CSV files only parse files with
given extension (default: csv)
-v, --verbose Be verbose
-s, --stats Print out Stats after running
-r, --redo Attempt to re-import a failed import or import more
data, uses stored metadata from previous run
--config-template-only
Configure the ElasticSearch template and then exit
--clear-interrupted-flag
Clear the interrupted flag, forcefully (NOT
RECOMMENDED)
-f INGEST_FILE, --file INGEST_FILE
Input CSV file
-d INGEST_DIRECTORY, --directory INGEST_DIRECTORY
Directory to recursively search for CSV files --
mutually exclusive to '-f' option
-D INGEST_DAY, --ingest-day INGEST_DAY
Day to use for metadata, in the format 'YYYY-MM-dd',
e.g., '2021-01-01'. Defaults to todays date, use
'YYYY-MM-00' to indicate a quarterly ingest, e.g.,
2021-04-00
-o COMMENT, --comment COMMENT
Comment to store with metadata
Performance Options:
--pipelines PIPELINES
Number of pipelines (default: 2)
--shipper-threads SHIPPER_THREADS
How many threads per pipeline to spawn to send bulk ES
messages. The larger your cluster, the more you can
increase this, defaults to 1
--fetcher-threads FETCHER_THREADS
How many threads to spawn to search ES. The larger
your cluster, the more you can increase this, defaults
to 2
--bulk-ship-size BULK_SHIP_SIZE
Size of Bulk Elasticsearch Requests (default: 10)
--bulk-fetch-size BULK_FETCH_SIZE
Number of documents to search for at a time (default:
50), note that this will be multiplied by the number
of indices you have, e.g., if you have 10
pydat-<number> indices it results in a request for 500
documents