
PyRIT
Open-source framework for red-teaming generative AI systems: automate attack prompts, score model responses, and audit behavior to identify security…

Open-source framework for red-teaming generative AI systems: automate attack prompts, score model responses, and audit behavior to identify security…

Modular LLM vulnerability scanner that probes for hallucination, data leakage, prompt injection, jailbreaks, and toxicity using static, dynamic, and…

LLM-driven agentic group shilling attack framework that manipulates black-box collaborative-filtering recommender rankings using adaptive multi-role…

A curated list of useful resources that cover Offensive AI.

An alignment auditing agent capable of quickly exploring alignment hypothesis

Interactive dashboards and libraries for responsible AI model debugging, covering error analysis, fairness, interpretability, counterfactuals, causal…

Research code for poisoning attacks on the PGM-index, demonstrating how to craft adversarial data to degrade learned index performance.

A diagnostic framework for measuring LLM vulnerability to Affective Contextual Erosion (ACE) and related liminal attack vectors. **Delirium** is not…

This repository hosts a multimodal web attack dataset (MWAD) to advance AI-driven threat detection research.

A comprehensive set of fairness metrics for datasets and machine learning models, explanations for these metrics, and algorithms to mitigate bias in…

Reverse Shell Detection with Machine Learning

A Binary Genetic Traits Lexer Framework

Voice-based detective interrogation game. Mistral Large 3 + Voxtral STT + ElevenLabs TTS. Built for the Mistral Worldwide Hackathon 2026.

Automated behavioral evaluation framework for LLMs that generates diverse test scenarios to probe for sycophancy, bias, and other safety-relevant…

Trajectory-aware evolutionary search framework for red-teaming LLM agents over MCP servers, generating adversarial prompts to map vulnerability…

image scaling attacks for multi-modal prompt injection

A novel adversarial attack on LLM based on the Exponentiated Gradient Descent technique.

Graph-based threat detection system using inexact graph vector matching to compare threat graphs with CTI-derived attack query graphs for automated…