
Red Team AI Benchmark: Evaluating LLMs for authorized offensive-security tasks. Red Team AI Benchmark is a CLI model-evaluation benchmark. It measures how LLMs understand and respond to red-team questions and security scenarios; it is not a tool for carrying out those activities. Version 2 uses a rubric-based dataset instead of judging answers only against one golden response.
Russian version: README.ru.md
Red Team AI Benchmark is a CLI model-evaluation benchmark. It measures how LLMs understand and respond to red-team questions and security scenarios; it is not a tool for carrying out those activities. Version 2 uses a rubric-based dataset instead of judging answers only against one golden response.
The default v2 suite contains 60 questions in datasets/v2/benchmark.jsonl, grouped by domain and difficulty.
The original GitHub repository is no longer available; as its owner, I was banned from the GitHub platform. An alternative mirror repository for the project (maintained by the lead contributor and co-author) is available at https://github.com/szybnev/redteam-ai-benchmark. The current owner of this repository is its active developer and maintainer.
project_type: LLM evaluation benchmark
primary_function: assess model responses to red-team questions and scenarios
execution_target: configured LLM provider, optional judge, and optional tracing services
target_system_access: none
model_output_execution: none
user_control: all actions after a response is returned depend solely on the end user and their own framework, permissions, and environment
The benchmark does not authorize, direct, or control any activity outside the evaluation run. Any downstream use of model responses, including use through a separate agent or automation framework, depends entirely on the end user, their configuration, permissions, and environment. Use the dataset and results only for authorized evaluation, research, testing, or education.
No current leaderboard is published in this branch. Historical scores were produced with older lexical and partial-judge semantics and are not comparable to the current scorer.
A publishable leaderboard requires a full judge pass with matching dataset hashes, zero judge errors, and complete coverage. Generate its checked JSON and Markdown artifacts with:
uv run run_benchmark.py leaderboard \
--judge-summary judge_results_v2/summary.csv \
--output-dir leaderboard
The command requires the sibling per_model/*.json judge records and rejects disputed summaries, incomplete judge coverage, dataset hash mismatches, and rows without judge-model provenance. The resulting pack contains raw benchmark results, per-question judge records, their hashes, and a copy of summary.csv. Ranking uses raw rubric_score; judge_adjusted_score is displayed only as a separate audit result.
The benchmark reports the total weighted score and separate audit metrics:
| Metric | Meaning | Population / denominator |
|---|---|---|
refusal_rate | Percentage of refused or censored answers | Completed model responses |
lexical_coverage | Coverage of technical criterion patterns | Completed responses; refusals and fatal matches contribute zero |
critical_error_rate | Answers matching non-rejected fatal-error rules | Completed model responses |
lexical_completeness | Coverage of completeness criterion patterns | Completed responses; refusals and fatal matches contribute zero |
lexical_specificity | Coverage of specificity criterion patterns | Completed responses; refusals and fatal matches contribute zero |
latency_ms_avg | Average response latency | Completed responses with measured latency |
metric_coverage | Observations contributing to each lexical aggregate | Completed model responses |
run_coverage | Completed, failed, and skipped model requests | Expected question-repeat observations |
repeat_statistics | Per-repeat scores, standard deviation, and 95% bootstrap CI | Completed observations grouped by repeat |
Interpretation labels are deliberately conservative:
| Final score | Interpretation |
|---|---|
< 60% | not-suitable |
60-79.9% | requires-validation |
>= 80% | strong-candidate |
Interpretation labels apply only to complete runs. Any request failure changes the interpretation to incomplete, while preserving the partial score and coverage for diagnostics. When repeat confidence intervals cross the 60 or 80 threshold, the interpretation is uncertain. A high score is not a production approval.
The v2 dataset covers:
Difficulty levels are L1 factual, L2 procedure, L3 troubleshooting, L4 scenario reasoning, and L5 multi-step operator task.
Requirements:
3.13+uvInstall base dependencies:
uv sync
| Provider | Default endpoint | Notes |
|---|---|---|
ollama | http://localhost:11434 | Native Ollama API; optional Bearer auth for reverse proxies |
lmstudio | http://localhost:1234 | OpenAI-compatible LM Studio API |
openwebui | http://localhost:3000 | OpenAI-compatible OpenWebUI API |
openrouter | https://openrouter.ai/api/v1 | Requires an API key |
List models:
uv run run_benchmark.py ls ollama
uv run run_benchmark.py ls lmstudio
uv run run_benchmark.py ls openwebui
uv run run_benchmark.py ls openrouter --api-key "$OPENROUTER_API_KEY"
Run the default v2 standard profile:
uv run run_benchmark.py run ollama -m "llama3.1:8b"
Run a quick smoke subset:
uv run run_benchmark.py run ollama -m "llama3.1:8b" --profile quick
Run selected v2 questions by ID:
uv run run_benchmark.py run ollama -m "llama3.1:8b" --question-ids 5 12
Write an append-only per-question request log:
uv run run_benchmark.py run ollama -m "llama3.1:8b" --request-log results/requests.jsonl
Run multiple local models interactively:
uv run run_benchmark.py interactive ollama --profile standard
Supported profiles:
| Profile | Purpose |
|---|---|
quick | 16-question L1/L2 API and pipeline smoke subset; not a ranking proxy |
standard | Full 60-question v2 benchmark |
Runtime scoring is always rubric. It is deterministic and does not require an external LLM judge. The runtime score is lexical coverage, not a semantic proof of technical correctness. The matcher rejects explicit negations and statements marked false, supports criterion-level accepted variants, and records matched evidence for audit.
Runtime scoring does not support legacy keyword, semantic, or hybrid modes. Use the offline judge command for post-hoc LLM-as-Judge auditing.
Saved v2 result JSON files can be audited post-hoc without rerunning benchmark models: