Skip to content
KitploitKITPLOIT
ToolsExploitsBlog
Log in
Submit
ToolsExploitsBlog
Submit

Hacking, PenTest, and Cybersecurity Tools for Your Security Arsenal!

Kitploit is a directory of hacking, cybersecurity, and pentesting tools. Discover the latest project updates to find vulnerabilities, analyze systems, automate testing, and strengthen your security.

FeedsContactPrivacy© 2026 Kitploit

Tool Directory

Categories

View all categories
Loading categories
benign-instruction-bench — Re-evaluating prompt-injection detectors on LLM agent tool outputs (paper draft, scripts, scores) | Kitploit
Tools/GitHubGitHub/lzwhehe/benign-instruction-bench
Defensive ToolsVulnerability AnalysisMachine LearningPapers & ResearchLearning & EducationAI Security
GitHublzwhehe/benign-instruction-bench

benign-instruction-bench

Re-evaluating prompt-injection detectors on LLM agent tool outputs (paper draft, scripts, scores)

View Repository
33 days agoNot yet reviewed

Most Popular

View all →

Discover the most used tools by our community.

Explore all tools

Browse our collection of tools

View all tools →
Share

Passing the Test You Trained On

Re-evaluating prompt-injection detectors where LLM agents actually use them: on the tool outputs an agent reads.

Teams pick injection detectors by their benchmark scores. We check whether those scores predict behavior inside an agent, and find that they mostly measure how close the benchmark is to the detector's training data.

Findings

Fifteen detectors (nine open, six license-gated including Meta's Prompt Guard 2) and two task-aware LLM judges, on two agent benchmarks (AgentDojo, tau-bench) and on BIPIA. Benign tool outputs come from replaying each benchmark's ground-truth tool calls without an LLM. Detection is TPR at the threshold that gives 1% FPR.

DetectorReleasedFPR on benign tool outputs (AgentDojo / tau-bench)Detection AgentDojoDetection tau-benchDetection BIPIA
Horizon-Labs guard-base20260.0% / 0.7%82.2%100.0%37.7%
Wolf Defender20260.9% / 0.0%73.8%90.8%3.5%
Prompt Guard 2 (86M)20250.6% / 0.0%69.4%57.9%10.4%
Prompt Guard 2 (22M)20250.0% / 0.0%60.9%60.8%31.7%
Prismor-1.5B20266.5% / 46.2%72.2%15.2%4.0%
Sheltron-68M202610.6% / 4.4%20.1%95.2%57.2%
Judge (Llama-3.1-8B)––55.3%78.3%8.3%
PIGuard202529.5% / 11.0%2.1%52.7%95.1%
ProtectAI v2202430.1% / 66.9%0.5%10.1%0.7%

(All 15 detectors and both judges: paper/tab_many.tex.)

  1. Detection rankings transfer poorly between benchmarks (Kendall τ over 15 detectors): 0.01 for BIPIA vs AgentDojo, 0.28 for AgentDojo vs tau-bench, 0.31 for BIPIA vs tau-bench; none significant. The BIPIA leader (PIGuard) ranks 13th of 15 on AgentDojo; Prismor drops from 72% on AgentDojo to 15% on tau-bench.
  2. False-positive rates on benign tool outputs range from 0% to over 90%, are consistent across the two agent benchmarks (τ = 0.67, p = 0.001), and are not predicted by false positives on ordinary text.
  3. The form of the training inputs, not overlap in attack strings, explains the results. PIGuard trained on 1,116 full BIPIA inputs and leads BIPIA. It also saw 53 of InjecAgent's 62 attacks, but only as short prompts; inside tau-bench tool outputs it detects seen and unseen InjecAgent attacks alike (52.1% vs 55.8%). Horizon-Labs shares no data with any benchmark, trained on agent-style inputs, and leads both agent benchmarks.
  4. Task-aware 7–8B judges (one trained before AgentDojo existed) reach 54–78% at 1% FPR on agent tool outputs, below the best dedicated detectors. Outside its training distribution every approach misses whole classes of injections; apart from PIGuard, no approach catches more than 39% of task-irrelevant injections.
  5. Removing serialization markup or declaring the input type lowers default-threshold false positives by up to about twentyfold but does not fix low-FPR detection.

All numbers are regenerated from the score files by analysis/compute_all.py; the paper reads them only through paper/numbers.tex.

Layout

scripts/     build evaluation sets, score detectors and judges
analysis/    compute_all.py, compute_many.py (every number, table, figure), make_macros.py (-> LaTeX macros)
results/     detector and judge scores used in the paper
paper/       LaTeX source, figures, compiled main.pdf

Reproduce

pip install -r requirements.txt
AGENTDOJO_PY=/path/to/agentdojo-env/bin/python QWEN=/models/Qwen2.5-7B-Instruct LLAMA=/models/Llama-3.1-8B-Instruct bash reproduce.sh

To regenerate only the numbers and figures from the released scores, copy results/*.json into the working directory, rebuild the .jsonl sets (steps 1–3 of reproduce.sh, CPU only), then run python analysis/compute_all.py && python analysis/make_macros.py.

Evaluation data is rebuilt from the public sources, not redistributed: AgentDojo v1.2, tau-bench (MIT), InjecAgent attacks (MIT), BIPIA (MIT), GitHub READMEs (h1alexbel/github-readmes), AESLC, FineWeb-Edu. Some scripts assume BIPIA and PIGuard are cloned under ~/agentsec/.

Building the paper

paper/usenix-2020-09-xetex.sty is a build copy of the USENIX style that compiles with tectonic/XeTeX. For submission, switch main.tex to the official usenix-2020-09.sty and compile with pdflatex.

Status

Draft. Not yet evaluated against adaptive attacks; both agent benchmarks are simulations and tau-bench's injection point is ours; commercial API detectors not included. See §6 of the paper.

License

Code, analysis scripts, and score files: MIT (see LICENSE). Third-party files keep their own terms: the USENIX LaTeX style (paper/usenix-2020-09*.sty) and the benchmarks and detectors the scripts download (AgentDojo, tau-bench, InjecAgent, BIPIA, and each detector's model license).

Download Tool