
promptfoo
Test your prompts, agents, and RAGs. Red teaming/pentesting/vulnerability scanning for AI. Compare performance of GPT, Claude, Gemini, DeepSeek, and…

Test your prompts, agents, and RAGs. Red teaming/pentesting/vulnerability scanning for AI. Compare performance of GPT, Claude, Gemini, DeepSeek, and…

Open-source framework for red-teaming generative AI systems: automate attack prompts, score model responses, and audit behavior to identify security…

Clusters and elements to attach to MISP events or attributes (like threat actors)

A repo for jailbreaking various LLMs, mainly Claude

Adversary Emulation Framework

A collection of awesome resources related AI security

Modular LLM vulnerability scanner that probes for hallucination, data leakage, prompt injection, jailbreaks, and toxicity using static, dynamic, and…

Fully automatic censorship removal for language models

LLM security testing framework for detecting prompt injection, jailbreaks, and adversarial attacks — 190+ probes, 28 providers, single Go binary

Security scanner for AI/ML model files. Detects malicious code, backdoors, and vulnerabilities before deployment

🐢 Open-Source Evaluation & Testing library for LLM Agents

A privacy-first app that strips AI watermarks from content you own.

Bypass llm guardrails by confusing it with fabricated tool output.

Open-source AI security platform providing perimeter defense for LLMs and AI agents through swarm analysis, policy enforcement, adversarial testing,…

Interactive dashboards and libraries for responsible AI model debugging, covering error analysis, fairness, interpretability, counterfactuals, causal…

:cloud: :zap: Granular, Actionable Adversary Emulation for the Cloud

Curated reading list and taxonomy of attack and defense research for mobile on-device AI systems, covering adversarial, backdoor, model stealing, and…

Research implementation for mitigating adaptive prompt injections via on-policy distillation, with training recipes and evaluators for SEP, PISmith,…