Skip to content
KitploitKITPLOIT
ToolsBlog
Log in
Submit
ToolsBlog
Submit

Hacking, PenTest, and Cybersecurity Tools for Your Security Arsenal!

Kitploit is a directory of hacking, cybersecurity, and pentesting tools. Discover the latest project updates to find vulnerabilities, analyze systems, automate testing, and strengthen your security.

··Feeds·Contact·Privacy·© 2026 Kitploit

Tool Directory

Categories

View all categories
Loading categories
Tools/GitLabGitLab/toxy4ny/redteam-ai-benchmark
Penetration TestingMachine LearningLearning & EducationRed TeamingAI SecurityLabs & Practice
GitLabtoxy4ny/redteam-ai-benchmark

redteam-ai-benchmark

View Repository
2212 months agoNot yet reviewed

Most Popular

View all →

Discover the most used tools by our community.

Explore all tools

Browse our collection of tools

View all tools →

About

Red Team AI Benchmark: Evaluating LLMs for authorized offensive-security tasks. Red Team AI Benchmark is a CLI model-evaluation benchmark. It measures how LLMs understand and respond to red-team questions and security scenarios; it is not a tool for carrying out those activities. Version 2 uses a rubric-based dataset instead of judging answers only against one golden response.

Share

Red Team AI Benchmark

Russian version: README.ru.md

Red Team AI Benchmark is a CLI model-evaluation benchmark. It measures how LLMs understand and respond to red-team questions and security scenarios; it is not a tool for carrying out those activities. Version 2 uses a rubric-based dataset instead of judging answers only against one golden response.

The default v2 suite contains 60 questions in datasets/v2/benchmark.jsonl, grouped by domain and difficulty.

Repository Status

The original GitHub repository is no longer available; as its owner, I was banned from the GitHub platform. An alternative mirror repository for the project (maintained by the lead contributor and co-author) is available at https://github.com/szybnev/redteam-ai-benchmark. The current owner of this repository is its active developer and maintainer.

Purpose and Scope

project_type: LLM evaluation benchmark
primary_function: assess model responses to red-team questions and scenarios
execution_target: configured LLM provider, optional judge, and optional tracing services
target_system_access: none
model_output_execution: none
user_control: all actions after a response is returned depend solely on the end user and their own framework, permissions, and environment

Explicit Non-goals

  • This repository is not a hacking tool, exploit framework, scanner, C2, persistence tool, payload runner, or autonomous red-team agent.
  • It does not discover, access, exploit, modify, or maintain access to target systems.
  • It does not execute model output. The benchmark only sends evaluation prompts to configured model endpoints, scores returned text, and writes results.
  • The presence of offensive-security topics in the test dataset describes the evaluation domain; it does not grant permission or provide authorization for activity against any system.

User Responsibility

The benchmark does not authorize, direct, or control any activity outside the evaluation run. Any downstream use of model responses, including use through a separate agent or automation framework, depends entirely on the end user, their configuration, permissions, and environment. Use the dataset and results only for authorized evaluation, research, testing, or education.

image

Published Leaderboard

No current leaderboard is published in this branch. Historical scores were produced with older lexical and partial-judge semantics and are not comparable to the current scorer.

A publishable leaderboard requires a full judge pass with matching dataset hashes, zero judge errors, and complete coverage. Generate its checked JSON and Markdown artifacts with:

uv run run_benchmark.py leaderboard \
  --judge-summary judge_results_v2/summary.csv \
  --output-dir leaderboard

The command requires the sibling per_model/*.json judge records and rejects disputed summaries, incomplete judge coverage, dataset hash mismatches, and rows without judge-model provenance. The resulting pack contains raw benchmark results, per-question judge records, their hashes, and a copy of summary.csv. Ranking uses raw rubric_score; judge_adjusted_score is displayed only as a separate audit result.

What v2 Measures

The benchmark reports the total weighted score and separate audit metrics:

MetricMeaningPopulation / denominator
refusal_ratePercentage of refused or censored answersCompleted model responses
lexical_coverageCoverage of technical criterion patternsCompleted responses; refusals and fatal matches contribute zero
critical_error_rateAnswers matching non-rejected fatal-error rulesCompleted model responses
lexical_completenessCoverage of completeness criterion patternsCompleted responses; refusals and fatal matches contribute zero
lexical_specificityCoverage of specificity criterion patternsCompleted responses; refusals and fatal matches contribute zero
latency_ms_avgAverage response latencyCompleted responses with measured latency
metric_coverageObservations contributing to each lexical aggregateCompleted model responses
run_coverageCompleted, failed, and skipped model requestsExpected question-repeat observations
repeat_statisticsPer-repeat scores, standard deviation, and 95% bootstrap CICompleted observations grouped by repeat

Interpretation labels are deliberately conservative:

Final scoreInterpretation
< 60%not-suitable
60-79.9%requires-validation
>= 80%strong-candidate

Interpretation labels apply only to complete runs. Any request failure changes the interpretation to incomplete, while preserving the partial score and coverage for diagnostics. When repeat confidence intervals cross the 60 or 80 threshold, the interpretation is uncertain. A high score is not a production approval.

Dataset Coverage

The v2 dataset covers:

  • Windows tradecraft
  • AD and AD CS
  • Web exploitation
  • Cloud and IAM
  • Containers and Kubernetes
  • Detection and evasion reasoning
  • OpSec and operational tradeoffs
  • Tool usage
  • Post-exploitation planning
  • Validation and reporting

Difficulty levels are L1 factual, L2 procedure, L3 troubleshooting, L4 scenario reasoning, and L5 multi-step operator task.

Installation

Requirements:

  • Python 3.13+
  • uv
  • One provider: Ollama, LM Studio, OpenWebUI, or OpenRouter

Install base dependencies:

uv sync

Providers

ProviderDefault endpointNotes
ollamahttp://localhost:11434Native Ollama API; optional Bearer auth for reverse proxies
lmstudiohttp://localhost:1234OpenAI-compatible LM Studio API
openwebuihttp://localhost:3000OpenAI-compatible OpenWebUI API
openrouterhttps://openrouter.ai/api/v1Requires an API key

Usage

List models:

uv run run_benchmark.py ls ollama
uv run run_benchmark.py ls lmstudio
uv run run_benchmark.py ls openwebui
uv run run_benchmark.py ls openrouter --api-key "$OPENROUTER_API_KEY"

Run the default v2 standard profile:

uv run run_benchmark.py run ollama -m "llama3.1:8b"

Run a quick smoke subset:

uv run run_benchmark.py run ollama -m "llama3.1:8b" --profile quick

Run selected v2 questions by ID:

uv run run_benchmark.py run ollama -m "llama3.1:8b" --question-ids 5 12

Write an append-only per-question request log:

uv run run_benchmark.py run ollama -m "llama3.1:8b" --request-log results/requests.jsonl

Run multiple local models interactively:

uv run run_benchmark.py interactive ollama --profile standard

Supported profiles:

ProfilePurpose
quick16-question L1/L2 API and pipeline smoke subset; not a ranking proxy
standardFull 60-question v2 benchmark

Scoring

Runtime scoring is always rubric. It is deterministic and does not require an external LLM judge. The runtime score is lexical coverage, not a semantic proof of technical correctness. The matcher rejects explicit negations and statements marked false, supports criterion-level accepted variants, and records matched evidence for audit.

Runtime scoring does not support legacy keyword, semantic, or hybrid modes. Use the offline judge command for post-hoc LLM-as-Judge auditing.

Offline LLM-as-Judge

Saved v2 result JSON files can be audited post-hoc without rerunning benchmark models:

Download Tool