
garak
Modular LLM vulnerability scanner that probes for hallucination, data leakage, prompt injection, jailbreaks, and toxicity using static, dynamic, and…

Modular LLM vulnerability scanner that probes for hallucination, data leakage, prompt injection, jailbreaks, and toxicity using static, dynamic, and…

Python library for adversarial machine learning security, enabling red and blue teams to run evasion, poisoning, extraction, and inference attacks…

Open-source framework for red-teaming generative AI systems: automate attack prompts, score model responses, and audit behavior to identify security…

Framework for auditing machine learning algorithms against adversarial attacks and biases, providing educational tools and an academic paper to…

LLM-driven agentic group shilling attack framework that manipulates black-box collaborative-filtering recommender rankings using adaptive multi-role…

An alignment auditing agent capable of quickly exploring alignment hypothesis

A curated list of useful resources that cover Offensive AI.

An adversarial example library for constructing attacks, building defenses, and benchmarking both

Research code for poisoning attacks on the PGM-index, demonstrating how to craft adversarial data to degrade learned index performance.

A Python toolbox to create adversarial examples that fool neural networks in PyTorch, TensorFlow, and JAX

Trajectory-aware evolutionary search framework for red-teaming LLM agents over MCP servers, generating adversarial prompts to map vulnerability…

Multi-layered prompt injection detector for AI applications using heuristics, LLM-based analysis, vectorDB attack signatures, and canary token leak…

Interactive dashboards and libraries for responsible AI model debugging, covering error analysis, fairness, interpretability, counterfactuals, causal…

Malware Mutation Using Reinforcement Learning and Generative Adversarial Networks

Reverse Shell Detection with Machine Learning

Voice-based detective interrogation game. Mistral Large 3 + Voxtral STT + ElevenLabs TTS. Built for the Mistral Worldwide Hackathon 2026.

Graph-based threat detection system using inexact graph vector matching to compare threat graphs with CTI-derived attack query graphs for automated…

A guided mutation-based fuzzer for ML-based Web Application Firewalls