
benign-instruction-bench
Re-evaluating prompt-injection detectors on LLM agent tool outputs (paper draft, scripts, scores)

Re-evaluating prompt-injection detectors on LLM agent tool outputs (paper draft, scripts, scores)

A Java Burp Plugin that performs text clustering on responses to identify outliers/groups based on the actual content of the server responses, say…

Runtime Application Self Protection for Python web servers, serverless functions and MCP servers, detecting attacks, prompt injection and data leaks…

Benchmark and defense code for persistent memory attacks on OpenClaw-style computer-use agents, with memory-zoning mitigation, attack scenarios, and…

A lifecycle benchmark for black-box LLM extraction attacks, defenses, and adaptive attacks.

Refusal localizes, the damage relocates, safety layers under few-sample fine-tuning

Research code for HARDE, an agent harness that probes and adaptively optimizes components for runtime risk detection and execution control across…

Benchmark harness measuring where prompt injection defenses fire in tool-using LLM agent pipelines, tracking canary tokens across exposed, persisted,…

AST-based Static Code Analyzer with Agentic LLM-Powered Relationship Mapping to discover Python RCE paths and deep deserialization chains on AI, LLM,…

Curated index of deep learning latency, energy-latency & timing (availability) attacks and defenses — companion to the ACM Computing Surveys…

Open-weight LLMs and training/evaluation code for defending against prompt injection attacks, with benchmarks for agentic tool-calling and…

Benchmark and evaluation harness testing whether LLM agents resist malicious instructions hidden in multimodal skill images, with 108 cases across…

Agentic jailbreak framework for LLM-based agents using scheme-based task decomposition, multi-turn disguising strategies, and adaptive self-evolution…

Evolutionary LLM jailbreak and guardrail framework that grows a reusable strategy pool via genetic mutation, Markov selection, and online adversarial…

Open framework for RL-based prompt injection red teaming, with a shared trainer, curriculum learning, and benchmarks like AgentDojo, InjecAgent, and…

Next-generation JavaScript identifier recovery with LLMs.

Adversarial image perturbation tool that uses SAM segmentation and CLIP models to evade AI-based scam image classifiers for security research.

Research code implementing T-Backdoor, temporal-trigger backdoor attacks on spiking neural networks using rate, latency, and jitter triggers without…