
PyRIT
Open-source framework for red-teaming generative AI systems: automate attack prompts, score model responses, and audit behavior to identify security…

Open-source framework for red-teaming generative AI systems: automate attack prompts, score model responses, and audit behavior to identify security…

Test your prompts, agents, and RAGs. Red teaming/pentesting/vulnerability scanning for AI. Compare performance of GPT, Claude, Gemini, DeepSeek, and…

Python library for adversarial machine learning security, enabling red and blue teams to run evasion, poisoning, extraction, and inference attacks…

Interactive dashboards and libraries for responsible AI model debugging, covering error analysis, fairness, interpretability, counterfactuals, causal…

Research toolkit for analyzing AI agent behavioral patterns through multi-disciplinary corpus analysis. Parses session logs, runs 23 analytical…

Security scanner for AI/ML model files. Detects malicious code, backdoors, and vulnerabilities before deployment

🐢 Open-Source Evaluation & Testing library for LLM Agents

Voice-based detective interrogation game. Mistral Large 3 + Voxtral STT + ElevenLabs TTS. Built for the Mistral Worldwide Hackathon 2026.

Fully automatic censorship removal for language models

A curated list of useful resources that cover Offensive AI.

Modular LLM vulnerability scanner that probes for hallucination, data leakage, prompt injection, jailbreaks, and toxicity using static, dynamic, and…

A privacy-first app that strips AI watermarks from content you own.

An alignment auditing agent capable of quickly exploring alignment hypothesis

Benchmark harness measuring where prompt injection defenses fire in tool-using LLM agent pipelines, tracking canary tokens across exposed, persisted,…

A lifecycle benchmark for black-box LLM extraction attacks, defenses, and adaptive attacks.

Benchmark and evaluation harness testing whether LLM agents resist malicious instructions hidden in multimodal skill images, with 108 cases across…

PISmith: Reinforcement Learning-based Red Teaming for Prompt Injection Defenses

Open framework for RL-based prompt injection red teaming, with a shared trainer, curriculum learning, and benchmarks like AgentDojo, InjecAgent, and…