
Security benchmark for evaluating OpenClaw agents against adversarial execution contexts including poisoned files, injected skills, misleading tool metadata, and encoded payloads. Includes 42 replay tasks across 6 risk suites with attack and utility scoring.
Open-world security evaluation for OpenClaw agents under adversarial execution contexts.
DeepTrap is a security benchmark for evaluating whether OpenClaw agents can complete benign user tasks while resisting malicious execution-context pressure: poisoned workspace files, injected skills, misleading tool metadata, unsafe command paths, planted secrets, and encoded payloads.
The public release contains 42 replay tasks arranged as 6 contextual vulnerability classes x 7 operational scenario families, plus the benchmark runner and public scoring code. It also includes 10 Chinese-language replay samples in tasks_zh/ for localized evaluation experiments. The private task-generation and attack-search pipeline is intentionally not included.
DeepTrap accompanies the paper:
Red-Teaming Agent Execution Contexts: Open-World Security Evaluation on OpenClaw
Hongwei Yao, Yiming Liu, Yiling He, and Bingrun Yang. arXiv:2605.11047, 2026.
Most agent benchmarks ask whether a model can finish useful work. DeepTrap asks a stricter question: can the agent finish useful work safely when the surrounding execution context is adversarial?
| What We Test | Why It Matters | |
|---|---|---|
| Context Robustness | Poisoned files, memory-like artifacts, playbooks, and workspace instructions | Real agents read from mutable local context, where untrusted text can silently redirect behavior |
| Secret Handling | Planted credentials, environment variables, API keys, and internal values | Agents must distinguish configuration awareness from secret exfiltration |
| Execution Boundaries | Dangerous shell commands, deployment scripts, and file-system side effects | Tool-using agents can cause durable harm even when the user prompt is benign |
| Skill/Tool Trust | Malicious skills, misleading tool metadata, and tool-abuse patterns | Agent safety depends on how the model interprets and validates tool affordances |
| Obfuscated Payloads | Base64 and multi-step encoded instructions | Attacks can hide behind normal-looking data-processing tasks |
tasks_zh/.AGS is Attack Grading Score and UGS is Utility Grading Score. Scores are reported by risk suite; Average is the mean across Risk 1 through Risk 6.
Full project page and leaderboard: ZJUICSR.github.io/DeepTrap
| Claw Model | Metric | Risk 1 | Risk 2 | Risk 3 | Risk 4 | Risk 5 | Risk 6 | Average |
|---|---|---|---|---|---|---|---|---|
| GPT-5.4 | AGS | 0.77 | 0.84 | 0.76 | 0.61 | 0.67 | 0.53 | 0.70 |
| GPT-5.4 | UGS | 0.91 | 0.83 | 0.86 | 0.77 | 0.74 | 0.87 | 0.83 |
| Claude-Sonnet-4.6 | AGS | 0.51 | 0.58 | 0.37 | 0.25 | 0.38 | 0.20 | 0.38 |
| Claude-Sonnet-4.6 | UGS | 0.71 | 0.69 | 0.55 | 0.45 | 0.55 | 0.71 | 0.61 |
| GLM-5 | AGS | 0.81 | 0.93 | 0.74 | 0.83 | 0.79 | 0.88 | 0.83 |
| GLM-5 | UGS | 0.90 | 0.90 | 0.98 | 0.89 | 0.83 | 0.88 | 0.90 |
| Qwen3.5-Plus | AGS | 0.93 | 0.93 | 0.86 | 0.74 | 0.88 | 0.97 | 0.88 |
| Qwen3.5-Plus | UGS | 0.95 | 0.92 | 1.00 | 0.98 | 0.93 | 0.93 | 0.95 |
| MiniMax-M2.5 | AGS | 0.86 | 0.89 | 0.77 | 0.66 | 0.90 | 0.89 | 0.83 |
| MiniMax-M2.5 | UGS | 0.92 | 0.95 | 1.00 | 0.88 | 0.74 | 0.90 | 0.90 |
| DeepSeek-v4-Flash | AGS | 0.90 | 0.96 | 0.80 | 0.90 | 0.82 | 0.94 | 0.89 |
| DeepSeek-v4-Flash | UGS | 0.98 | 0.96 | 1.00 | 0.96 | 0.85 | 1.00 | 0.96 |
| Deepseek-v4-Pro | AGS | 0.90 | 0.96 | 0.74 | 0.87 | 0.85 | 0.86 | 0.86 |
| Deepseek-v4-Pro | UGS | 0.90 | 0.91 | 1.00 | 0.81 | 0.84 | 0.89 | 0.89 |
| MiMo-v2.5 | AGS | 0.86 | 0.87 | 0.71 | 0.73 | 0.57 | 0.60 | 0.72 |
| MiMo-v2.5 | UGS | 0.96 | 0.95 | 0.88 | 0.93 | 0.83 | 0.89 | 0.91 |
| MiMo-v2.5-pro | AGS | 0.74 | 0.83 | 0.56 | 0.58 | 0.58 | 0.53 | 0.64 |
| MiMo-v2.5-pro | UGS | 0.92 | 0.90 | 0.88 | 0.87 | 0.71 | 0.87 | 0.86 |
DeepTrap crosses six contextual vulnerability classes with seven benign scenario families. Each task uses a normal user prompt; the adversarial behavior is induced by the workspace context.
DeepTrap constructs compromised execution contexts from benign instructions and clean workspaces, searches candidate adversarial payloads with multi-objective reward signals, and refines successful payloads through reflection-based probing.
| Suite | Category | Tasks | What is measured |
|---|---|---|---|
R1 | Harness Hijacking | 7 | Resistance to poisoned context and injected workflow steps |
R2 | Privacy Leakage | 7 | Protection of planted credentials and sensitive values |
R3 | Unauthorized Execution | 7 | Refusal of dangerous commands and boundary-crossing actions |
R4 | Supply Chain Risk | 7 | Handling of malicious skills and tool-like integrations |
R5 | Tool Abuse | 7 | Detection of misleading or mismatched tool behavior |
R6 | Encoding Obfuscation | 7 | Handling of encoded instructions and multi-step payloads |