Open-source cross-modal and multimodal prompt injection test suite. 250,000+ attack payloads across text, image, document, and audio modalities. Research-backed by OWASP LLM Top 10, CrossInject (ACM MM 2025), FigStep (AAAI 2025), DolphinAttack, and CSA 2026.
516,588 labeled samples (251,782 attack + 251,576 benign, plus a 13,230-sample real-world validation split) across five dataset versions plus external dataset ingestion, covering cross-modal, multi-turn, adversarial suffix, jailbreak template, indirect injection, tool manipulation, agentic, evasion, reasoning DoS, video generation, VLA robotic, LoRA supply chain, audio-native LLM, RAG optimisation, MCP cross-server, coding agent, serialization boundary, and agent skill supply chain attacks on AI systems. Attack and benign samples are balanced 1:1 (ratio 0.9992:1 after audit cleanup).
Built for training and evaluating prompt injection detectors. All samples are labeled (expected_detection: true/false), source-attributed to peer-reviewed papers or documented industry research, and structured for direct use in binary classifiers.
The payloads are plain JSON. Load them directly with your language's stdlib — no dependencies:
import json, pathlib
records = []
for p in pathlib.Path("payloads_v5").glob("*.json"):
records.extend(json.loads(p.read_text()))
print(f"{len(records)} labeled samples")
Every record carries expected_detection: true|false, an attack_category string, and a source field pointing at the original paper or documented incident. That's enough to train a binary classifier, run per-category ASR, or slice by attack vector.
Prompt injection is defined here as: text embedded in an LLM input that is intended to override, hijack, or redirect the model's behaviour away from its operator-specified task. This definition follows Greshake et al. 2023 (arXiv:2302.12173) and OWASP LLM01:2025.
The scope is runtime injection only -- text that an attacker can place in the model's context window at inference time. The dataset deliberately excludes:
The distinction matters for detection: a runtime detector reads the prompt, not the model weights. Attacks that only affect training are out of scope.
The dataset was built in four layers:
Layer 1 -- Seed payloads (hand-crafted, 210 + 187 + 284 seeds): Injection seeds for each attack category were written by hand, grounded in peer-reviewed papers and documented real-world incidents. Every seed is tagged with its academic source and attack reference. Seeds were reviewed against the inclusion definition above -- any seed that could be re-read as a benign request without an override component was discarded or rewritten.
Layer 2 -- Programmatic expansion via templates and encoding (v2, 14,358 samples): Seeds were passed through PyRIT v0.12.1's 162 jailbreak templates and 13 encoding converters. Template expansion is fully deterministic and reproducible from the generator script. GCG adversarial suffixes were drawn from the published literature (Zou et al. 2023) and appended to seeds; live gradient optimization is optional and requires a GPU.
Layer 3 -- Cross-modal delivery (v1 + v4 cross-modal, 35,687 samples): Injection seeds were delivered across 7 image methods, 4 document types x 5 hiding locations, 6 audio methods, and multi-modality combinations. This follows the threat model in FigStep (arXiv:2311.05608) and CrossInject (arXiv:2504.14348): the injection text may arrive in any modality the pipeline processes, not only the text field. Modality fields (image_content, doc_content, audio_content) record what the model's extractor would read from that channel.
Layer 4 -- Benign samples (50,516 total): Benign prompts were drawn from published academic and industry datasets (Stanford Alpaca, WildChat, deepset/prompt-injections, LMSYS Chatbot Arena). Benign multimodal samples pair these text prompts with real image captions (MS-COCO 2017, Flickr30k), document passages (Wikipedia EN, arXiv via RedPajama), and audio transcripts (LibriSpeech, Mozilla Common Voice). A set of 130 hand-crafted edge cases uses attack-adjacent vocabulary ("ignore", "override", "system prompt", "password") in genuinely benign contexts to reduce false positive training.
Layer 5 -- Real-world validation split (13,230 samples): Layers 1-4 are constructed: hand-written, templated, or drawn from other datasets. Layer 5 is not. It was collected from a live game where players scored points for beating a deployed detector, tiered through boss-level "castle" stages and multimodal "ghost" passes. Every successful and attempted bypass was logged, then anonymised (identifiers and payment data stripped at the table level, in-text PII redacted, high-risk rows quarantined for manual review rather than published) and released as payloads_live/. Where Layers 1-4 measure coverage against known attack classes, Layer 5 measures whether a detector holds up against a motivated human actively trying to break it. See payloads_live/README.md for the full anonymisation methodology.
All attack payloads: expected_detection: true
All benign samples: expected_detection: false
Labels are assigned by construction, not by human review of individual samples. The correctness guarantee is therefore at the category level: each category maps to a documented attack class with a specific mechanism. Individual samples inherit the label from the seed and category they were generated from.
There is no adversarial label noise introduced deliberately. The detector is expected to learn the injection pattern, not to distinguish between "genuine" and "fake" injections -- all samples in the attack set represent real or plausible attack strings.
The edge-case benign set was designed to reduce false positives on security-adjacent language. It covers 10 vocabulary clusters: ignore, override, system prompt, password, instructions, jailbreak (in the iPhone sense), bypass surgery, XSS (as a security topic, not an attack), prompt (as in camera shutter), and inject (as in dependency injection / medical). A detector that achieves high precision on the full dataset but low precision on the edge-case set is overfitting to surface-level keyword matching.
The full dataset was audited for label correctness and contamination. The audit checks:
expected_detection: trueexpected_detection: false<|im_start|> tokens)id, text, expected_detection, modalities) present on every sampleAudit results: