
benign-instruction-bench
Re-evaluating prompt-injection detectors on LLM agent tool outputs (paper draft, scripts, scores)

Re-evaluating prompt-injection detectors on LLM agent tool outputs (paper draft, scripts, scores)

A lifecycle benchmark for black-box LLM extraction attacks, defenses, and adaptive attacks.

Benchmark harness measuring where prompt injection defenses fire in tool-using LLM agent pipelines, tracking canary tokens across exposed, persisted,…

Open framework for RL-based prompt injection red teaming, with a shared trainer, curriculum learning, and benchmarks like AgentDojo, InjecAgent, and…

Research code implementing T-Backdoor, temporal-trigger backdoor attacks on spiking neural networks using rate, latency, and jitter triggers without…

finance domain specific red-teaming benchmark evaluation rubric

Benchmarking prompt injection detections for web agents.

Defensive framework that maintains a safety-focused shadow memory to detect and block prompt-injection and long-horizon threats against LLM agents…

Proof-of-concept and research repository for CVE-2024-37054, an unsafe deserialization flaw in MLflow PyFunc model loading that can lead to remote…

Research code for red-teaming AI auto-mode monitors, including simulation evals, fuzzing, and monitor implementations for Claude Code and Codex…

Syntactic Ghost: An Imperceptible General-purpose Backdoor Attacks on Pre-trained Language Models

Research code reproducing multi-turn LLM jailbreak experiments (FITD, MRCJ, ActorAttack, X-Teaming) from the SoK intent-oriented systematization…

Research code for Rubric-Induced Preference Drift (RIPD): evolutionary rubric search, benchmark-preserving selection, and DPO policy misalignment…

Multi-agent automated context management for long horizon tasks in local AI Agents

This repository contains the official implementation of the paper "[Safety in Batches? Understanding and Mitigating Safety Failures in Batch…

LLM-driven agentic group shilling attack framework that manipulates black-box collaborative-filtering recommender rankings using adaptive multi-role…

Hybrid machine-learning pipelines for detecting SQL injection in web traffic, combining DistilBERT and BERT-GNN models with adversarial training and…

Research code for extracting and training safety-awareness directions in multimodal LLMs to improve refusal behavior while limiting benign-task drift.