
Evaluation framework for AI penetration testing agents that measures validated vulnerability discovery using LLM-based semantic matching, bipartite resolution, and cumulative analysis across real-world targets.
AI pentesting agents are increasingly credible as offensive security systems, but current benchmarks still provide limited guidance on which systems will perform best on real-world targets. Most existing evaluations assess and optimize for predefined goals such as flag capture, remote code execution, exploit reproduction, or trajectory similarity, in simplified or narrow settings. These benchmarks are valuable for measuring bounded capabilities, yet they do not adequately capture the complexity, open-ended exploration, and strategic decision-making required in realistic pentesting. We present a practical evaluation framework that shifts assessment from task completion to validated vulnerability discovery, allowing evaluation in sufficiently complex targets spanning multiple attack surfaces and vulnerability classes. The framework combines structured ground-truth with LLM-based semantic matching to identify vulnerabilities, bipartite resolution to score findings under realistic ambiguity, continuous ground-truth maintenance, repeated and cumulative evaluation of stochastic agents, efficiency metrics, and reduced-suite selection for sustainable experimentation. This methodology extends the state of the art by enabling more realistic and operationally informative comparison of AI pentesting agents. To enable reproducibility, we additionally release expert-annotated ground-truth and code for the proposed evaluation protocol.
Evaluation pipeline for security testing tools. Compares tool findings against ground truth datasets using LLM-based matching and produces precision, recall, F1, and F0.5 metrics.
poetry install
Requires Python 3.11+ and Poetry installed.
# 1. Set your LLM API key
export OPENAI_API_KEY="..."
# 2. Run evaluation
ethibench evaluate ./my_experiment --dataset path/to/dataset.yaml
# 3. View results
cat ./my_experiment/evaluation_outputs/summary.md
ethibench evaluateRuns the full evaluation pipeline on an experiment directory.
ethibench evaluate <experiment_dir> --dataset <dataset.yaml> [options]
# Batch: evaluate all experiments in a folder
ethibench evaluate --parent-dir final_experiments/ --dataset <dataset.yaml>
# Force re-evaluation (ignore cached artifacts)
ethibench evaluate <experiment_dir> --dataset <dataset.yaml> --force
Arguments:
experiment_dir — (optional) Directory containing target subdirectories (or run_* subdirectories each with target subdirs). Can be omitted when using --parent-dir.Options:
--dataset, -d — (required) Path to dataset YAML file.--gt-dir, -g — Ground truth directory. Defaults to gt/ next to the dataset YAML.--output-dir, -o — Output directory. Defaults to evaluation_outputs/ inside experiment dir. Ignored in batch mode.--replicates, -n — Number of LLM matching replicates (default: 1).--force, -f — Re-run all steps, ignoring cached artifacts. By default, existing intermediate results (raw matchings, bipartite matchings, metrics) are reused.--parent-dir, -p — Parent folder containing multiple experiment directories to evaluate in batch. All immediate subdirectories are treated as experiments.What it does:
target_id), loads each findings.jsonl, assigns subset_name from dataset YAML.metrics.json if present, aggregates cost/token/duration.evaluation_outputs/plots/.evaluation_outputs/summary.md.ethibench analyzeRuns analysis tools on existing evaluation outputs.
ethibench analyze <experiment_dir> --dataset <dataset.yaml> [options]
# Batch: analyze all experiments and produce aggregated results
ethibench analyze --parent-dir final_experiments/ --dataset <dataset.yaml>
Arguments:
experiment_dir — (optional) Experiment directory to analyze. Can be omitted when using --parent-dir.Options:
--dataset, -d — (required) Path to dataset YAML file.--gt-dir, -g — Ground truth directory. Defaults to gt/ next to the dataset YAML.--output-dir, -o — Evaluation outputs directory. Defaults to evaluation_outputs/ inside experiment dir.--parent-dir, -p — Parent folder containing multiple experiment directories to analyze in batch. Produces per-experiment analysis plus aggregated results.Per-experiment outputs (evaluation_outputs/analysis/):
duplicates.json — findings matched in raw matching but removed by bipartite optimization.unmatched.json — findings with no ground truth match (false positives).statistics.json — GT coverage stats, findings-per-GT distribution.Aggregated outputs (only with --parent-dir, in <parent-dir>/aggregated_analysis/):
all_duplicates.jsonl — all duplicate findings across all experiments (JSONL, full finding objects with experiment field).all_false_positives.jsonl — all unmatched/false positive findings across all experiments (JSONL format).gt_statistics_avg.json — averaged GT coverage per subset, plus per-experiment coverage summary.ethibench compareCompares evaluation results across multiple experiments, producing side-by-side plots and a summary report. Each experiment must already have evaluation outputs (run ethibench evaluate first). Labels are always the directory names.
# Explicit experiment directories
ethibench compare exp-gpt4o/ exp-claude/ --output-dir comparison/
# Auto-discover all experiments under a parent folder
ethibench compare --parent-dir all-experiments/ --output-dir comparison/
# Mix: explicit dirs + auto-discovery
ethibench compare exp-extra/ --parent-dir all-experiments/ --output-dir comparison/
Arguments:
experiment_dirs — (optional) One or more experiment directories to include explicitly.Options:
--output-dir, -o — (required) Output directory for comparison results.--parent-dir, -p — Parent folder to auto-discover experiments from. Any immediate subdirectory containing an evaluation_outputs/ folder is included, sorted alphabetically. Can be combined with explicit experiment_dirs.Outputs (in --output-dir):
comparison.json — raw comparison data for all experiments.plots/ — side-by-side PNG charts.comparison.md — Markdown summary.pairwise_comparison.md — pairwise A/B statistical comparison (top 4 experiments by F1).pairwise_comparison.tex — LaTeX version of the pairwise table.cumulative-analysis/ — (if cumulative data exists) delta analysis comparing averaged vs cumulative F1, plus cumulative comparison plots.findings.jsonl)One JSON object per line. Required field: title, description. Optional: url, cwe, severity, score, steps, evidence, metadata, etc.
Each findings.jsonl lives inside a target directory — the directory name determines which target the findings belong to.
{"title": "SQL Injection in Login", "description": "User input not sanitized", "cwe": "89"}
*_gt.jsonl)One JSON object per line.
{"id": "gt-001", "name": "SQL Injection", "subset_name": "MyApp", "target_id": "app", "category": "CWE-89", "description": "Database query vulnerability", "cvss": 9.8}
- subset: "MyApp"
weight: 1.0
targets:
- target_id: "app"