
StepJack
Benchmarking framework for evaluating computer-use AI agents against multi-step indirect prompt injection, with automatic adversarial goal…

Benchmarking framework for evaluating computer-use AI agents against multi-step indirect prompt injection, with automatic adversarial goal…

LLM-driven agentic group shilling attack framework that manipulates black-box collaborative-filtering recommender rankings using adaptive multi-role…

Black-box attack framework that hijacks reasoning in agentic retrieval-augmented generation systems by injecting poisoned documents, with support for…

A diagnostic framework for measuring LLM vulnerability to Affective Contextual Erosion (ACE) and related liminal attack vectors. **Delirium** is not…

Defense framework that keeps audio-language models frozen and adds a mid-layer risk gate with late-layer safety adapters to block audio jailbreaks at…

Metasploit for machine learning.

PISmith: Reinforcement Learning-based Red Teaming for Prompt Injection Defenses

Black-box input-stage purification defense that neutralizes backdoor attacks on object detectors via corruption, diffusion reconstruction, and DBSCAN…