
promptfoo
Test your prompts, agents, and RAGs. Red teaming/pentesting/vulnerability scanning for AI. Compare performance of GPT, Claude, Gemini, DeepSeek, and…

Test your prompts, agents, and RAGs. Red teaming/pentesting/vulnerability scanning for AI. Compare performance of GPT, Claude, Gemini, DeepSeek, and…

Adversary Emulation Framework

Bypass llm guardrails by confusing it with fabricated tool output.

CVE-2026-6765, Test only FormAutofill handlers exposed in Firefox

Evades LLM text watermarks by injecting Unicode variation selectors; includes the SynthID generator, mean-g detector, normalization defenses, and…

Detection-aware BloodHound attack-path scoring - the quietest route to your objective, calibrated across five detection tiers…

Self-Defeating Audits: reproducible lab showing a low-privilege PostgreSQL role reversibly blinding a trigger-based auditor + poisoning attribution…

Detects LLM context-leakage attacks by training lightweight behavior probes on log-probabilities, with vLLM offline/server detection pipelines.

The TTPForge is a Cybersecurity Framework for developing, automating, and executing attacker Tactics, Techniques, and Procedures (TTPs).

CVE-2026-20685 - Draft or TODO

Reproduces CVE-2026-21019 by manipulating node clock to force early Kubernetes CronJob execution; includes vulnerable YAML manifest and Python…

Proof-of-concept exploit for CVE-2026-21003 demonstrating JWT authentication bypass by omitting the kid header and using the 'none' algorithm to…

Browser PoC demonstrating CVE-2026-2828, a WebGPU timing side-channel that leaks cross-origin iframe pixel values by measuring GPU timestamp-query…

Advanced CVE-2023-44487 HTTP/2 Rapid Reset vulnerability exploitation framework. Features multi-connection concurrent attacks, adaptive rate control,…

Open-source cross-modal and multimodal prompt injection test suite. 250,000+ attack payloads across text, image, document, and audio modalities.…

Black-box attack framework that hijacks reasoning in agentic retrieval-augmented generation systems by injecting poisoned documents, with support for…

This is the tool to dump the LSASS process on modern Windows 11

Benchmark for evaluating AI agent safety against attacks embedded in skill-facing context, with 155 cases across 6 risk domains, measuring task…