
FARO - Document Sensitivity Detector

FARO is a tool for detecting sensitive information in documents in an organization. It is oriented to be used by small companies and particulars that want to track their sensitive documents inside their organization but who cannot spend much time and money configuring complex Data Protection tools.
FARO extract sensitivity indicators from documents (e.g. Document IDs, monetary quantities, personal emails) and gives a sensitivity score to the document (from low to high) using the frequence and type of the indicators in the document.
Currently all the functionality of this tool is for documents written in Spanish although it can be easily upgraded to cover more languages.
This tool is developed by TEGRA R&D Cybersecurity Center.
The project contains the following folders:
faro/ : this is FARO module with the main functionality.config/: yaml configuration files go here. There is one yaml file per language (plus one nolanguage.yaml to provide basic functionality for non detected languages) and one yaml file with common configurations for all languages config/commons.yaml.models/: this is the folder to place FARO models.faro_detection.py: launcher of FARO for standalone operation over a single file.faro_spider.sh: script for bulk processing.docker_build_faro.sh: script for building FARO docker image on Linux and Mac OS.docker_build_faro.bat: script for building FARO docker image on Windows.docker_run_faro.sh: script for running a FARO container on Linux and Mac OS.docker_run_faro.bat: script for running a FARO container on Windows.CHANGELOG: FARO changelog.FARO can run as a standalone container using Docker. You can build the image for yourself or get it from Docker Hub repository.
Provided you have Docker installed and running on your system, execute the following command to get the latest FARO image from Docker Hub.
docker pull gradiant/faro
To run the docker image use the scripts docker_run_faro.sh (Linux/Mac OS) or docker_run_faro.bat (Windows). You can find them in the root of the project or in the latest release.
Provided you have Docker installed and running on your system, do the following to build the FARO image.
Linux and Mac OS
./docker_build_faro.sh
Windows
docker_build_faro.bat
To run a FARO container some scripts are provided in the root of the project. You can copy and use those scripts from anywhere else for convenience. The "output" folder will be created under your current directory.
Linux and Mac OS
./docker_run_faro.sh <your folder with files>
Windows
docker_run_faro.bat <your folder with files>
We have added OCR support to tika through its tesseract integration. Some customization of the OCR process can be tweaked through the use of an env file whose path needs to be provided as the second argument to the script. We have provided a commented example to serve as a template here
./docker_run_faro.sh <your folder with files> <path to env file>
for example:
./docker_run_faro.sh ../data docker_faro_env_example.list
FARO creates an "output" folder inside the current folder and stores the results of the execution in two files:
output/scan.$CURRENT_TIME.csv: is a csv file with the score given to the document and the frequence of indicators in each file.filepath,score,person_position_organization,monetary_quantity,signature,personal_email,mobile_phone_number,financial_data,document_id,custom_words,meta:content-type,meta:author,meta:pages,meta:lang,meta:date,meta:filesize,meta:num_words,meta:num_chars,meta:ocr
/Users/test/code/FARO_datasets/quick_test_data/Factura_NRU_0_1_001.pdf,high,0,0,0,0,0,0,1,4,application/pdf,Powered By Crystal,1,es,,85739,219,1185,False
/Users/test/code/FARO_datasets/quick_test_data/Factura_Plancha.pdf,high,0,6,0,0,0,0,2,8,application/pdf,Python PDF Library - http://pybrary.net/pyPdf/,1,es,,77171,259,1524,True
/Users/test/code/FARO_datasets/quick_test_data/20190912-FS2019.pdf,high,0,3,0,0,0,0,1,2,application/pdf,FPDF 1.6,1,es,2019-09-12T20:08:19Z,1545,62,648,False
output/scan.$CURRENT_TIME.entity: is a json with the list of indicators (disaggregated) extracted in a file. For example:{"filepath": "/Users/test/code/FARO_datasets/quick_test_data/Factura_NRU_0_1_001.pdf", "entities": {"custom_words": {"facturar": 3, "total": 1}, "prob_currency": {"12,0021": 1, "12,00": 1, "9,92": 1, "3,9921": 1, "3,99": 1, "3,30": 1, "15,99": 1, "13,21": 1, "1.106.166": 1, "1,00": 1, "99,00": 1}, "document_id": {"89821284M": 1}}, "datetime": "2019-12-11 14:19:17"}
{"filepath": "/Users/test/code/FARO_datasets/quick_test_data/Factura_Plancha.pdf", "entities": {"document_id": {"H82547761": 1, "21809943D": 2}, "custom_words": {"factura": 2, "facturar": 2, "total": 2, "importe": 2}, "monetary_quantity": {"156,20": 4, "2,84": 2, "0,00": 2, "159,04": 2, "32,80": 4, "191,84": 2}, "prob_currency": {"1,00": 6, "189,00": 2}}, "datetime": "2019-12-11 14:19:27"}
{"filepath": "/Users/test/code/FARO_datasets/quick_test_data/20190912-FS2019.pdf", "entities": {"document_id": {"C-01107564": 1}, "custom_words": {"factura": 1, "total": 1}, "monetary_quantity": {"3,06": 1, "0,64": 1, "3,70": 1}}, "datetime": "2019-12-11 14:19:33"}
NOTE: ONLY LINUX AND MAC OS X
The mode requires some operative system and libraries in order to work properly
It is advisable to use a separate virtual environment. To instantiate a virtual environment with virtualenv.
virtualenv -p `which python3` <yourenvname>
To activate the virtual environment on your terminal just type:
source <yourenvname>/bin/activate
The easiest way of getting the system up and running is to install the dependencies this way
pip install -r requirements.txt
The list of dependencies are the following:
These other dependencies are used for testing:
FARO relies on several ML models in order to work.