
PyRIT
Open-source framework for red-teaming generative AI systems: automate attack prompts, score model responses, and audit behavior to identify security…

Open-source framework for red-teaming generative AI systems: automate attack prompts, score model responses, and audit behavior to identify security…

A diagnostic framework for measuring LLM vulnerability to Affective Contextual Erosion (ACE) and related liminal attack vectors. **Delirium** is not…

A productionized greedy coordinate gradient (GCG) attack tool for large language models (LLMs)

Malware Mutation Using Reinforcement Learning and Generative Adversarial Networks

Evaluation framework that tests whether large language models follow invisible Unicode-encoded instructions embedded in normal-looking text, with…

Automated behavioral evaluation framework for LLMs that generates diverse test scenarios to probe for sycophancy, bias, and other safety-relevant…

A Binary Genetic Traits Lexer Framework

Framework for auditing machine learning algorithms against adversarial attacks and biases, providing educational tools and an academic paper to…

Trajectory-aware evolutionary search framework for red-teaming LLM agents over MCP servers, generating adversarial prompts to map vulnerability…

image scaling attacks for multi-modal prompt injection

Black-box attack framework that hijacks reasoning in agentic retrieval-augmented generation systems by injecting poisoned documents, with support for…

A novel adversarial attack on LLM based on the Exponentiated Gradient Descent technique.

This repository hosts a multimodal web attack dataset (MWAD) to advance AI-driven threat detection research.

Graph-based threat detection system using inexact graph vector matching to compare threat graphs with CTI-derived attack query graphs for automated…

Robust audio watermarking framework embedding binary messages into magnitude spectrograms, with differentiable attack simulation, Q-Former pooling,…

Multi-layered prompt injection detector for AI applications using heuristics, LLM-based analysis, vectorDB attack signatures, and canary token leak…

Simulates CVE-2024-38063 TCP/IP remote code execution attack, captures network traffic with TShark, and trains a machine learning model to detect…

Extracts structured Cyber Threat Intelligence (CTI) from PDF, DOCX, and TXT documents using LLMs. Generates MITRE ATT&CK attack flows, detection…