
The full repo of all the labs available as part of the benchmark
A benchmark for evaluating AI agents on web security challenges, generated by the TarantuLabs engine.
TarantuBench is a collection of 100 vulnerable web applications, each containing a hidden flag (TARANTU{...}). An agent's job is to find and extract the flag by interacting with the application over HTTP — just like a human pentester would.
The challenges range from beginner-level SQL injection login bypasses to advanced multi-step attack chains that require exploiting up to 5 vulnerabilities in sequence — including business logic abuse, stored XSS for session theft, JWT forgery, SSRF, and SQL injection on internal APIs.
Every lab is a self-contained Node.js/Express application with an in-memory SQLite database. No external dependencies, no network access needed — just boot the server and start probing.
All challenges in this release were generated using TarantuLabs' proprietary lab generation engine.
TarantuBench v1 represents a mature, scalable benchmark backed by a proven generation pipeline:
Node harness requirements: Node.js 18+ and npm.
Inspect AI task requirements: Python 3.11+, Docker, and uv or another
PEP 517-compatible installer.
The runnable lab dataset is published on Hugging Face at
tarantulabs/TarantuBench.
This GitHub repository contains the evaluation harness and documentation.
git clone https://github.com/Trivulzianus/TarantuBench.git
cd TarantuBench
cd eval && npm install && cd ..
# Download the dataset file from Hugging Face, or clone the dataset repo:
# git clone https://huggingface.co/datasets/tarantulabs/TarantuBench data
# Run your agent against all 100 labs
node eval/harness.js --dataset data/tarantubench-v1.jsonl \
--command "python my_agent.py --url {URL}" \
--timeout 300
# Generate scorecard
node eval/scorecard.js
Before running a formal evaluation, validate that the local or Hugging Face dataset has the expected row count and schema:
node eval/validate-dataset.js --dataset data/tarantubench-v1.jsonl --expected-count 100
node eval/validate-dataset.js --hf tarantulabs/TarantuBench --expected-count 100
The harness boots each lab, places a transparent logging proxy in front of it, and runs your agent command (replacing {URL} with the target address). Your agent can be written in any language — it just needs to make HTTP requests and submit the flag via POST {URL}/submit-flag with body {"flag": "TARANTU{...}"}.
# Boot one lab in server mode — harness prints the URL, you connect your agent
node eval/harness.js --dataset data/tarantubench-v1.jsonl \
--labs corporate-portal-chain-xss-idor \
--mode server --timeout 300
Each row in data/tarantubench-v1.jsonl represents one challenge:
| Column | Type | Description |
|---|---|---|
lab_id | string | Unique identifier |
title | string | Human-readable challenge name |
description | string | Brief scenario description (shown to the agent) |
objectives | list[string] | What the agent is told to accomplish |
hints | list[string] | Optional progressive hints (for ablation studies) |
difficulty | string | Beginner, Intermediate, or Advanced |
category | string | Primary vulnerability family (e.g., SQL Injection, XSS) |
vuln_subtype | string | Specific technique (e.g., sqli-union, xss-stored) |
chain_type | string or null | Multi-step chain ID, or null for single-vulnerability labs |
server_code | string | Full Node.js/Express source code for the vulnerable application |
dependencies | object | npm package dependencies needed to run the server |
| Difficulty | Count | Description |
|---|---|---|
| Beginner | 35 | Single vulnerability, direct exploitation |
| Intermediate | 25 | Requires enumeration, filter bypass, or multi-step logic |
| Advanced | 40 | Multi-step chains, business logic flaws, or deep exploitation |
| Category | Count |
|---|---|
| Multi-Vulnerability Chains | 34 |
| SQL Injection | 20 |
| IDOR (Insecure Direct Object Reference) | 11 |
| Auth/Authz Bypass | 10 |
| XSS (Cross-Site Scripting) | 10 |
| Business Logic | 8 |
| Command Injection | 5 |
| SSRF | 2 |
34 of the 100 labs require chaining multiple vulnerabilities:
| Chain Type | Count | Steps |
|---|---|---|
| SSRF → SQL Injection | 8 | Bypass access control via SSRF, then extract flag via SQLi |
| SSRF → Blind SQLi | 5 | SSRF to reach internal endpoint, then blind boolean extraction |
| XSS → SQL Injection | 7 | Steal admin session via stored XSS, then use admin-only search with SQLi |
| XSS → IDOR | 5 | Steal admin session via stored XSS, then access hidden data via IDOR |
| JWT Forgery → Blind SQLi | 4 | Crack weak JWT secret, forge elevated token, extract flag char-by-char |
| JWT Forgery → IDOR | 3 | Crack JWT, forge elevated role, access restricted API endpoints |
| Biz Logic → XSS → JWT → SSRF → SQLi | 1 | 5-step chain through referral abuse, session theft, JWT forgery, SSRF pivot, and union SQLi |
| XSS → JWT → SSRF → SQLi | 1 | 4-step chain through session theft, JWT forgery, SSRF, and SQL injection |
Labs are distributed across 20 realistic application themes — banking portals, hospital systems, e-commerce stores, IoT dashboards, government services, gaming platforms, and more — ensuring vulnerability patterns are tested in diverse contexts.
TarantuBench also exposes an Inspect AI task for
the inspect_evals beta registry flow. The task keeps the lab dataset on
Hugging Face, boots each generated Node/Express app inside an Inspect Docker
sandbox, and gives the model configurable constrained tools rather than a shell.
uv sync
uv run inspect eval src/tarantubench/task.py@tarantubench \
--model openai/gpt-4o \
--limit 1