
RLCDAlignBench
Benchmark suite and code for detecting AI alignment failures, with 44 benchmarks across ten failure types and a zero-shot RLCD detector evaluated on…

Benchmark suite and code for detecting AI alignment failures, with 44 benchmarks across ten failure types and a zero-shot RLCD detector evaluated on…

finance domain specific red-teaming benchmark evaluation rubric

Black-box input-stage purification defense that neutralizes backdoor attacks on object detectors via corruption, diffusion reconstruction, and DBSCAN…

Benchmarking prompt injection detections for web agents.

Defensive framework that maintains a safety-focused shadow memory to detect and block prompt-injection and long-horizon threats against LLM agents…

Measuring Open Privilege in Agent Defenses

Experiments for control-token chain-of-thought suppression and parser-leniency attacks on tool-using LLM agents

Proof-of-concept that poisons MLflow registered models via the REST API, embedding a malicious pickle to trigger RCE when the model is loaded.

Proof-of-concept and research repository for CVE-2024-37054, an unsafe deserialization flaw in MLflow PyFunc model loading that can lead to remote…

Research code for red-teaming AI auto-mode monitors, including simulation evals, fuzzing, and monitor implementations for Claude Code and Codex…

Bidirectional token-classification model for PII detection and masking in text, with CLI for redaction, evaluation, and finetuning on-premises.

Research code for a gray-box trojan attack that flips a single KV-cache bit in fine-tuned LLM classifiers and measures per-class attack success rate.

Pre-execution action-auditing defense that detects and masks indirect prompt injection in tool-using LLM agents using embedding retrieval and…

Fix-Like Artifacts With Embedded Defects

Research code and experiments for defending tool-integrated LLM agents against adversarial attacks, extending Agent Security Bench with new defense…

Syntactic Ghost: An Imperceptible General-purpose Backdoor Attacks on Pre-trained Language Models

Research code reproducing multi-turn LLM jailbreak experiments (FITD, MRCJ, ActorAttack, X-Teaming) from the SoK intent-oriented systematization…

Research code for Rubric-Induced Preference Drift (RIPD): evolutionary rubric search, benchmark-preserving selection, and DPO policy misalignment…