
CVE-Factory
CVE-Factory is a Multi-Agent system for fully automated, end-to-end CVE reproduction. Given CVE records, the system automatically researches details, generates test cases, builds Docker environments, and validates that each vulnerability can be both exploited and patched. The pipeline transforms CVE metadata into reproducible, testable vulnerability environments without manual intervention.
⚠️ Security Warning: This system builds and runs Docker containers containing vulnerable software. You MUST use the Docker-in-Docker (DinD) environment to isolate CVE containers from your host system. Never run CVE-Factory directly on your host Docker daemon.
Input CVE records, get a complete CVE reproduction environment. Following the Terminal Bench standard, each generated task package includes:
Dockerfile and docker-compose.yaml hosting the vulnerable applicationtask.yaml containing structured instruction descriptions (CVE-identity-free)solution.sh to patch the vulnerabilityrun-tests.sh to start the evaluationSpecifically designed for security tasks, our testing logic is split into:
No manual research, no manual coding - fully automated from raw CVE metadata to validated reproduction.
Generated Artifact Structure:
CVE-2025-XXXX/
├── task.yaml # Structured Task Metadata
├── Dockerfile # Vulnerable Environment Setup
├── docker-compose.yaml # Service Orchestration
├── task-deps/
├── solution.sh # Verified Patch
└── test/
├── test_func.py # Functionality Check
├── test_vuln.py # Vulnerability Exploit Check
└── run-tests.sh # One-click Evaluation Script
In a large-scale evaluation of 554 CVEs from 2025, CVE-Factory successfully reproduced 499 cases, achieving an 90.1% success rate. Furthermore, a rigorous expert review of 471 successful cases confirmed that 312 tasks (66.2%) were completely and accurately reproduced!
When compared against security experts using identical initial information, our system achieved a ~95% verification pass rate on environment and solution construction — demonstrating expert-level capability in automated vulnerability reproduction.
📂 Open Dataset: We release 1,000+ CVE task environments in the
cve_tasks/directory:
trainset/(887 tasks): Used for training Abacus-cve. The 4,000+ distilled agent traces on Hugging Face 🤗 are generated from these tasks using Claude Opus 4.5 with a Mini SWE-Agent harness.trainset-2/: Additional tasks with relatively simpler difficulty. Not included in the training data.- NEW: Additional 3,181 tasks available at
cve_tasks_3k_compressedon Hugging Face (compressed archive due to size limits), with 18.8k agent traces for training Abacus-cve-v1.1.
Fine-tuning on CVE-Factory traces yields dramatic improvements across security benchmarks. Qwen3-32B achieves ~6.8× improvement on LiveCVEBench (5.29% → 35.79%), ~4.2× on PatchEval (5.66% → 23.58%), and even shows significant gains on Terminal-Bench (12.50% → 28.75%) — demonstrating strong cross-task generalization.
| Model | LiveCVEBench | PatchEval | Terminal-Bench | Avg |
|---|---|---|---|---|
| Qwen3-32B (base) | 5.29 | 5.66 | 12.50 | 7.82 |
| Abacus-cve (Ours) | 35.79 | 23.58 | 28.75 | 29.37 |
| Qwen3-Coder-30B | 10.58 | 9.91 | 13.75 | 11.41 |
| Qwen3-Coder-480B | 19.58 | 19.34 | 36.25 | 25.06 |
| MiniMax-M2 | 24.87 | 19.34 | 37.50 | 27.24 |
| Claude Sonnet 4 | 20.11 | 22.64 | 33.75 | 25.50 |
| Claude Sonnet 4.5 | 34.39 | 28.77 | 45.00 | 36.05 |
| Claude Opus 4.5 | 41.27 | 32.08 | 48.75 | 40.70 |
With just 4k traces, Abacus-cve (32B) outperforms Qwen3-Coder-480B, MiniMax-M2, and Claude Sonnet 4, approaching Claude Sonnet 4.5 level on security tasks.
NEW: Abacus-cve-v1.1 trained on 18.8k traces achieves further gains (+3.83 on LiveCVEBench, +2.38 on PatchEval). See cve_train_v1.1 for the expanded training data.
Unlike rigid retrieval workflows or simple tool-use loops, each agent operates as a full Claude Code session. We do not hard-code steps; instead, we define each agent by its Role (e.g., Analyzer), Goal (e.g., "Build a vulnerable environment"), Resources (e.g., Access to specific docs), and Verification Method (e.g., "Must pass check_env_ready"). Agents act like human developers: they autonomously explore files, debug errors, read logs, and iterate on solutions within their designated workspace.
CVE-Factory is designed to handle multiple CVEs simultaneously. Each CVE pipeline executes asynchronously, meaning faster tasks proceed to subsequent stages without waiting for slower ones. The system uses an asynchronous architecture that allows you to separate concurrency limits for each specific Agent type. For example, you can set a higher limit for lightweight research tasks (Analyzer) and a lower limit for resource-intensive Docker tasks (Builder). This flexibility prevents system overload while maximizing processing speed. Stage-level timeouts ensure that hung processes don't block the processing queue.