
T-MAP
Trajectory-aware evolutionary search framework for red-teaming LLM agents over MCP servers, generating adversarial prompts to map vulnerability…

Trajectory-aware evolutionary search framework for red-teaming LLM agents over MCP servers, generating adversarial prompts to map vulnerability…

Black-box attack framework that hijacks reasoning in agentic retrieval-augmented generation systems by injecting poisoned documents, with support for…

Research implementation of a poisoning attack against retrieval-augmented language models using camouflaged documents to evade filtering defenses and…

Benchmark for evaluating AI agent safety against attacks embedded in skill-facing context, with 155 cases across 6 risk domains, measuring task…

Bypass llm guardrails by confusing it with fabricated tool output.

Multi-stage prompt injection technique that bypasses LLM safety alignment via identity reassignment, refusal suppression, and output coercion,…

Research implementation for mitigating adaptive prompt injections via on-policy distillation, with training recipes and evaluators for SEP, PISmith,…

Deterministic memory-poisoning / prompt-injection measurement axis — CoSnitch (CVE-2026-24301) anchored. Inspect scorer, signed receipts.…

Python library for adversarial machine learning security, enabling red and blue teams to run evasion, poisoning, extraction, and inference attacks…

An adversarial example library for constructing attacks, building defenses, and benchmarking both

A Python toolbox to create adversarial examples that fool neural networks in PyTorch, TensorFlow, and JAX

Protection against Model Serialization Attacks

Algorithms for outlier, adversarial and drift detection

A security scanner for your LLM agentic workflows

Curated list of backdoor learning papers, surveys, and toolboxes, organizing poisoning-based attacks and defenses in deep learning for researchers…

Train, evaluate, and explore neural networks with built-in adversarial robustness tools, including PGD attacks, adversarial training, and robust…

Standardized adversarial robustness benchmark with a public leaderboard and downloadable model zoo for evaluating ML models against Lp attacks and…

Interpretability and explainability of data and machine learning models