
Veritensor v1.9.2
The Anti-Virus for AI Artifacts & RAG Firewall. A static analysis tool scanning Models and Notebooks for RCE, Datasets and RAG docs for Data Poisoning, PII, and Prompt Injections. Secure your AI Supply Chain.
🛡️ Veritensor: AI Data & Artifact Security
Veritensor is the Anti-Virus for AI Artifacts and the ultimate Firewall for RAG pipelines. It secures the entire AI Supply Chain by scanning the artifacts that traditional SAST tools miss: Models, Datasets, RAG Documents, and Notebooks.
Veritensor shift security left. Instead of waiting for a prompt injection to hit your LLM, Veritensor intercepts and sanitizes malicious documents, poisoned datasets, and compromised dependencies before they enter your Vector DB or execution environment.
Unlike standard SAST tools (which focus on code), Veritensor understands the binary and serialized formats used in Machine Learning:
- Models: Deep AST analysis of Pickle, PyTorch, Keras, Safetensors to block RCE and backdoors.
- Data & RAG: Streaming scan of Parquet, CSV, Excel, PDF to detect Data Poisoning, Prompt Injections, and PII.
- Notebooks: Hardening of Jupyter (.ipynb) files by detecting leaked secrets (using Entropy analysis), malicious magics, and XSS.
- Supply Chain: Audits dependencies (
requirements.txt,poetry.lock) for Typosquatting and known CVEs (via OSV.dev). - Agentic AI & MCP Servers: Pure AST analysis of Python files to detect Agent Hijacking risks in
@mcp.tool()functions. Scansclaude_desktop_config.jsonandmcp.jsonfor over-privileged permissions. - Governance: Generates cryptographic Data Manifests (Provenance) and signs containers via Sigstore.
🚀 Features
- Native RAG Security: Embed Veritensor directly into
LangChain,LlamaIndex,ChromaDB, andUnstructured.ioto block threats at runtime. - High-Performance Parallel Scanning: Utilizes all CPU cores with robust SQLite Caching (WAL mode). Re-scanning a 100GB dataset takes milliseconds if files haven't changed.
- Advanced Stealth Detection: Hackers hide prompt injections using CSS (
font-size: 0,color: white) and HTML comments. Veritensor scans raw binary streams to catch what standard parsers miss. - Dataset Security: Streams massive datasets (100GB+) to find "Poisoning" patterns (e.g., "Ignore previous instructions") and malicious URLs in Parquet, CSV, JSONL, and Excel.
- Archive Inspection: Safely scans inside .zip, .tar.gz, .whl files without extracting them to disk (Zip Bomb protected).
- Dependency Audit: Checks
pyproject.toml,poetry.lock, andPipfile.lockfor malicious packages (Typosquatting) and vulnerabilities. - Data Provenance: Command
veritensor manifest .creates a signed JSON snapshot of your data artifacts for compliance (EU AI Act). - Identity Verification: Automatically verifies model hashes against the official Hugging Face registry to detect Man-in-the-Middle attacks.
- De-obfuscation Engine: Automatically detects and decodes Base64 strings to uncover hidden payloads (e.g.,
SWdub3Jl...->Ignore previous instructions). - Magic Number Validation: Detects malware masquerading as safe files (e.g., an
.exerenamed toinvoice.pdf). - Smart Filtering & Entropy Analysis: Drastically reduces false positives in Jupyter Notebooks. Uses Shannon Entropy to find real, unknown API keys (WandB, Pinecone, Telegram) while ignoring safe UUIDs and standard imports.
- CISO-Ready HTML Reports: Generate beautiful, standalone HTML security reports with interactive charts and severity breakdowns using the
--htmlflag.
📦 Installation
Veritensor is modular. Install only what you need to keep your environment lightweight (~50MB core).
| Option | Command | Use Case |
|---|---|---|
| Core | pip install veritensor | Base scanner (Models, Notebooks, Dependencies) |
| RAG | pip install "veritensor[rag]" | Documents (PDF, DOCX, PPTX) |
| PII | pip install "veritensor[pii]" | ML-based PII detection (Presidio) |
| AWS | pip install "veritensor[aws]" | Direct scanning from S3 buckets |
| All | pip install "veritensor[all]" | Full suite for enterprise security |
Via Docker (Recommended for CI/CD)
docker pull arseniibrazhnyk/veritensor:latest
⚡ Quick Start
1. Scan a local project (Parallel)
Recursively scan a directory for all supported threats using 4 CPU cores:
veritensor scan ./my-rag-project --recursive --jobs 4
2. Scan RAG Documents & Excel
Check for Prompt Injections and Formula Injections in business data:
veritensor scan ./finance_data.xlsx
veritensor scan ./docs/contract.pdf
3. Generate Data Manifest
Create a compliance snapshot of your dataset folder:
veritensor manifest ./data --output provenance.json
4. Sync Policy-as-Code to Control Plane
Push your local security thresholds to the Enterprise Server:
veritensor scan . --sync-policy --api-key "vt_your_key"
5. Scan AI Datasets (Bias & Poisoning)
Veritensor uses streaming to handle huge files. It samples 10k rows by default for speed.
veritensor scan ./data/train.parquet --full-scan
6. Verify Model Integrity
Ensure the file on your disk matches the official version from Hugging Face (detects tampering):
veritensor scan ./pytorch_model.bin --repo meta-llama/Llama-2-7b
7. Scan from Amazon S3
Scan remote assets without manual downloading:
veritensor scan s3://my-ml-bucket/models/llama-3.pkl
8. Verify against Hugging Face
Ensure the file on your disk matches the official version from the registry (detects tampering):
veritensor scan ./pytorch_model.bin --repo meta-llama/Llama-2-7b
9. License Compliance Check
Veritensor automatically reads metadata from safetensors and GGUF files. If a model has a Non-Commercial license (e.g., cc-by-nc-4.0), it will raise a HIGH severity alert.
To override this (Break-glass mode), use:
veritensor scan ./model.safetensors --force
10. Scan AI Datasets
Veritensor uses streaming to handle huge files. It samples 10k rows by default for speed.
veritensor scan ./data/train.parquet --full-scan
11. Scan Jupyter Notebooks
Check code cells, markdown, and saved outputs for threats:
veritensor scan ./research/experiment.ipynb
12. Generate a CISO-Friendly HTML Report
Create a standalone, interactive HTML dashboard of your scan results:
veritensor scan ./project --html
13. Scan MCP Servers for Agent Hijacking
Detect dangerous tool logic and over-privileged permissions in your AI agent infrastructure:
# Scan MCP server Python files (AST analysis — no code execution)
veritensor scan ./mcp_servers/
# Also audit MCP configuration files
veritensor scan ./claude_desktop_config.json