
pgm-attack
Research code for poisoning attacks on the PGM-index, demonstrating how to craft adversarial data to degrade learned index performance.

Research code for poisoning attacks on the PGM-index, demonstrating how to craft adversarial data to degrade learned index performance.

Black-box attack framework that hijacks reasoning in agentic retrieval-augmented generation systems by injecting poisoned documents, with support for…

Research implementation of a poisoning attack against retrieval-augmented language models using camouflaged documents to evade filtering defenses and…

Benchmark for evaluating AI agent safety against attacks embedded in skill-facing context, with 155 cases across 6 risk domains, measuring task…

Research implementation for mitigating adaptive prompt injections via on-policy distillation, with training recipes and evaluators for SEP, PISmith,…

A Python toolbox to create adversarial examples that fool neural networks in PyTorch, TensorFlow, and JAX

Protection against Model Serialization Attacks

Train, evaluate, and explore neural networks with built-in adversarial robustness tools, including PGD attacks, adversarial training, and robust…

Threadless Process Injection using remote function hooking.

C# Reflective loader for unmanaged binaries.

Interpretability and explainability of data and machine learning models

:cloud: :zap: Granular, Actionable Adversary Emulation for the Cloud

Open-source framework for red-teaming generative AI systems: automate attack prompts, score model responses, and audit behavior to identify security…

Interactive dashboards and libraries for responsible AI model debugging, covering error analysis, fairness, interpretability, counterfactuals, causal…

A comprehensive set of fairness metrics for datasets and machine learning models, explanations for these metrics, and algorithms to mitigate bias in…

Detects LLM context-leakage attacks by training lightweight behavior probes on log-probabilities, with vLLM offline/server detection pipelines.

PoC script for HTTP/2 Rapid Reset (CVE-2023-44487) that sends crafted HTTP/2 streams to trigger denial-of-service conditions on vulnerable servers,…

Proof-of-concept exploit for CVE-2026-73292: CSRF attack on Semaphore UI password change endpoint, serving a malicious page that silently resets an…