Skip to content
KitploitKITPLOIT
ToolsExploitsBlog
Log in
Submit
ToolsExploitsBlog
Submit

Hacking, PenTest, and Cybersecurity Tools for Your Security Arsenal!

Kitploit is a directory of hacking, cybersecurity, and pentesting tools. Discover the latest project updates to find vulnerabilities, analyze systems, automate testing, and strengthen your security.

FeedsContactPrivacy© 2026 Kitploit

Tool Directory

Categories

View all categories
Loading categories
TarantuBench — The full repo of all the labs available as part of the benchmark | Kitploit
Tools/GitHubGitHub/trivulzianus/tarantubench
Authentication & AuthorizationVulnerability ScannersWeb Application ExploitationWeb SecurityCTFPenetration TestingLearning & EducationPayload DevelopmentLabs & Practice
GitHubtrivulzianus/tarantubench

TarantuBench

The full repo of all the labs available as part of the benchmark

21134 months agoNot yet reviewed

Most Popular

View all →

Discover the most used tools by our community.

Explore all tools

Browse our collection of tools

View all tools →
Share
View Repository

TarantuBench v1

A benchmark for evaluating AI agents on web security challenges, generated by the TarantuLabs engine.

What is this?

TarantuBench is a collection of 100 vulnerable web applications, each containing a hidden flag (TARANTU{...}). An agent's job is to find and extract the flag by interacting with the application over HTTP — just like a human pentester would.

The challenges range from beginner-level SQL injection login bypasses to advanced multi-step attack chains that require exploiting up to 5 vulnerabilities in sequence — including business logic abuse, stored XSS for session theft, JWT forgery, SSRF, and SQL injection on internal APIs.

Every lab is a self-contained Node.js/Express application with an in-memory SQLite database. No external dependencies, no network access needed — just boot the server and start probing.

All challenges in this release were generated using TarantuLabs' proprietary lab generation engine.

v1 — Generation at Scale

TarantuBench v1 represents a mature, scalable benchmark backed by a proven generation pipeline:

  • Throughput. The pipeline generates approximately 100 verified labs per hour using Claude Opus with adaptive thinking. Each lab is a complete, themed web application with realistic UI, seeded data, and one or more exploitable vulnerabilities.
  • Verification. Every generated lab is deterministically validated: boot the server, run an automatically generated solver, and confirm the flag is extractable. The pipeline achieves a 93% first-pass verification rate. Failed labs are automatically diagnosed and regenerated until the full batch passes.
  • Node.js/Express by design. All labs target Node.js/Express — this is a deliberate choice, not a limitation. It enables every challenge to run interactively in the browser via WebContainers on tarantulabs.com, making the benchmark accessible without any local setup.
  • What's next. Future versions will expand the vulnerability infrastructure to additional server frameworks and languages, and explore security challenges beyond web applications — including binary exploitation, network security, and cryptographic attacks.

Quick Start

Node harness requirements: Node.js 18+ and npm.

Inspect AI task requirements: Python 3.11+, Docker, and uv or another PEP 517-compatible installer.

The runnable lab dataset is published on Hugging Face at tarantulabs/TarantuBench. This GitHub repository contains the evaluation harness and documentation.

git clone https://github.com/Trivulzianus/TarantuBench.git
cd TarantuBench
cd eval && npm install && cd ..

# Download the dataset file from Hugging Face, or clone the dataset repo:
# git clone https://huggingface.co/datasets/tarantulabs/TarantuBench data

# Run your agent against all 100 labs
node eval/harness.js --dataset data/tarantubench-v1.jsonl \
  --command "python my_agent.py --url {URL}" \
  --timeout 300

# Generate scorecard
node eval/scorecard.js

Before running a formal evaluation, validate that the local or Hugging Face dataset has the expected row count and schema:

node eval/validate-dataset.js --dataset data/tarantubench-v1.jsonl --expected-count 100
node eval/validate-dataset.js --hf tarantulabs/TarantuBench --expected-count 100

The harness boots each lab, places a transparent logging proxy in front of it, and runs your agent command (replacing {URL} with the target address). Your agent can be written in any language — it just needs to make HTTP requests and submit the flag via POST {URL}/submit-flag with body {"flag": "TARANTU{...}"}.

Run a Single Lab Manually

# Boot one lab in server mode — harness prints the URL, you connect your agent
node eval/harness.js --dataset data/tarantubench-v1.jsonl \
  --labs corporate-portal-chain-xss-idor \
  --mode server --timeout 300

Why this benchmark?

  • Unambiguous evaluation: Either the agent extracts the flag, or it doesn't. No partial credit, no human judgment needed.
  • Graduated difficulty: From beginner (single injection point) to advanced (multi-step exploit chains requiring 2–5 chained vulnerabilities).
  • Reproducible: Deterministic server code and no external APIs. The harness generates a fresh per-run flag, so exploit behavior is reproducible while the literal flag value changes each run.
  • Generated at scale: Labs are synthesized by the TarantuLabs engine, not hand-written. This means the benchmark can grow programmatically as we add new vulnerability types and chain definitions.

Dataset Schema

Each row in data/tarantubench-v1.jsonl represents one challenge:

ColumnTypeDescription
lab_idstringUnique identifier
titlestringHuman-readable challenge name
descriptionstringBrief scenario description (shown to the agent)
objectiveslist[string]What the agent is told to accomplish
hintslist[string]Optional progressive hints (for ablation studies)
difficultystringBeginner, Intermediate, or Advanced
categorystringPrimary vulnerability family (e.g., SQL Injection, XSS)
vuln_subtypestringSpecific technique (e.g., sqli-union, xss-stored)
chain_typestring or nullMulti-step chain ID, or null for single-vulnerability labs
server_codestringFull Node.js/Express source code for the vulnerable application
dependenciesobjectnpm package dependencies needed to run the server

Challenge Breakdown

By Difficulty

DifficultyCountDescription
Beginner35Single vulnerability, direct exploitation
Intermediate25Requires enumeration, filter bypass, or multi-step logic
Advanced40Multi-step chains, business logic flaws, or deep exploitation

By Category

CategoryCount
Multi-Vulnerability Chains34
SQL Injection20
IDOR (Insecure Direct Object Reference)11
Auth/Authz Bypass10
XSS (Cross-Site Scripting)10
Business Logic8
Command Injection5
SSRF2

Chain Challenges

34 of the 100 labs require chaining multiple vulnerabilities:

Chain TypeCountSteps
SSRF → SQL Injection8Bypass access control via SSRF, then extract flag via SQLi
SSRF → Blind SQLi5SSRF to reach internal endpoint, then blind boolean extraction
XSS → SQL Injection7Steal admin session via stored XSS, then use admin-only search with SQLi
XSS → IDOR5Steal admin session via stored XSS, then access hidden data via IDOR
JWT Forgery → Blind SQLi4Crack weak JWT secret, forge elevated token, extract flag char-by-char
JWT Forgery → IDOR3Crack JWT, forge elevated role, access restricted API endpoints
Biz Logic → XSS → JWT → SSRF → SQLi15-step chain through referral abuse, session theft, JWT forgery, SSRF pivot, and union SQLi
XSS → JWT → SSRF → SQLi14-step chain through session theft, JWT forgery, SSRF, and SQL injection

Application Themes

Labs are distributed across 20 realistic application themes — banking portals, hospital systems, e-commerce stores, IoT dashboards, government services, gaming platforms, and more — ensuring vulnerability patterns are tested in diverse contexts.

Evaluation Harness

Inspect AI Task

TarantuBench also exposes an Inspect AI task for the inspect_evals beta registry flow. The task keeps the lab dataset on Hugging Face, boots each generated Node/Express app inside an Inspect Docker sandbox, and gives the model configurable constrained tools rather than a shell.

uv sync
uv run inspect eval src/tarantubench/task.py@tarantubench \
  --model openai/gpt-4o \
  --limit 1
Download Tool