
PyRIT
Open-source framework for red-teaming generative AI systems: automate attack prompts, score model responses, and audit behavior to identify security…

Open-source framework for red-teaming generative AI systems: automate attack prompts, score model responses, and audit behavior to identify security…

Clusters and elements to attach to MISP events or attributes (like threat actors)

Security scanner for AI/ML model files. Detects malicious code, backdoors, and vulnerabilities before deployment

The AI Security Verification Standard (AISVS) focuses on providing developers, architects, and security professionals with a structured checklist to…

An alignment auditing agent capable of quickly exploring alignment hypothesis

Tracker of publicly reported prompt-injection techniques, broken down by delivery method, encoding, and propagation behavior, with confirmed models,…

Benchmark harness measuring where prompt injection defenses fire in tool-using LLM agent pipelines, tracking canary tokens across exposed, persisted,…

A lifecycle benchmark for black-box LLM extraction attacks, defenses, and adaptive attacks.

Benchmark and evaluation harness testing whether LLM agents resist malicious instructions hidden in multimodal skill images, with 108 cases across…

Open framework for RL-based prompt injection red teaming, with a shared trainer, curriculum learning, and benchmarks like AgentDojo, InjecAgent, and…

Set of tools to assess and improve LLM security.

ExploitGym is a large-scale, realistic benchmark built from real-world vulnerabilities designed to evaluate AI agents' ability to develop exploits.

Evolutionary LLM jailbreak and guardrail framework that grows a reusable strategy pool via genetic mutation, Markov selection, and online adversarial…

Refusal localizes, the damage relocates, safety layers under few-sample fine-tuning

Defense framework that keeps audio-language models frozen and adds a mid-layer risk gate with late-layer safety adapters to block audio jailbreaks at…

Research code for a gray-box trojan attack that flips a single KV-cache bit in fine-tuned LLM classifiers and measures per-class attack success rate.

Experiments for control-token chain-of-thought suppression and parser-leniency attacks on tool-using LLM agents

An autonomous red-teaming engine for LLMs. RedThread manages the full security lifecycle: generating adversarial attacks, executing precision…