
bloom
Automated behavioral evaluation framework for LLMs that generates diverse test scenarios to probe for sycophancy, bias, and other safety-relevant…

Automated behavioral evaluation framework for LLMs that generates diverse test scenarios to probe for sycophancy, bias, and other safety-relevant…

Evaluation framework that tests whether large language models follow invisible Unicode-encoded instructions embedded in normal-looking text, with…

LLM security testing framework for detecting prompt injection, jailbreaks, and adversarial attacks — 190+ probes, 28 providers, single Go binary

The Python Risk Identification Tool for generative AI (PyRIT) is an open source framework built to empower security professionals and engineers to…

Open-source framework for red-teaming generative AI systems: automate attack prompts, score model responses, and audit behavior to identify security…

MAAD Attack Framework - An attack tool for simple, fast & effective security testing of M365 & Entra ID (Azure AD).

Adversary Emulation Framework

Framework for auditing machine learning algorithms against adversarial attacks and biases, providing educational tools and an academic paper to…

A collection of awesome resources related AI security

Purple Team Exercise Framework

Scripted framework for simulating over 50 MITRE ATT&CK techniques to test blue team detection capabilities. Includes Python scripts and a compiled…

Playbook-based adversary simulation framework that compiles JSON-defined attack paths into position-independent shellcode payloads for validating…

StealthRL: RL framework for adversarially paraphrasing AI text to stress-test detector robustness.

Attack framework for breaking fine-tuning based prompt injection defenses (SecAlign, SecAlign++, StruQ) using architecture-aware adversarial attacks…


Trajectory-aware evolutionary search framework for red-teaming LLM agents over MCP servers, generating adversarial prompts to map vulnerability…

Implementation of paper "DeeCLIP: A Robust and Generalizable Transformer-Based Framework for Detecting AI-Generated Images"

Automated framework for hijacking safety reasoning in large reasoning models via simulated reasoning traces and iterative prompt refinement to…