Skip to content
KitploitKITPLOIT
ToolsExploitsBlog
Log in
Submit
ToolsExploitsBlog
Submit

Hacking, PenTest, and Cybersecurity Tools for Your Security Arsenal!

Kitploit is a directory of hacking, cybersecurity, and pentesting tools. Discover the latest project updates to find vulnerabilities, analyze systems, automate testing, and strengthen your security.

··Feeds·Contact·Privacy·© 2026 Kitploit

Tool Directory

Categories

View all categories
Loading categories
DeepTrap — Security benchmark for evaluating OpenClaw agents against adversarial execution contexts including poisoned files, injected skills, misleading tool metadata, and encoded payloads. Includes 42 replay tasks across 6 risk suites with attack and utility scoring. | Kitploit
Tools/GitHubGitHub/zjuicsr/deeptrap
Privilege EscalationData ExfiltrationPenetration TestingCommand and ControlSupply Chain SecurityLearning & EducationRed TeamingPayload DevelopmentAI SecurityAdversarial AttackLabs & Practice
101164 months agoNot yet reviewed

Most Popular

View all →

Discover the most used tools by our community.

Explore all tools

Browse our collection of tools

View all tools →
Share
GitHub
zjuicsr/deeptrap

DeepTrap

Security benchmark for evaluating OpenClaw agents against adversarial execution contexts including poisoned files, injected skills, misleading tool metadata, and encoded payloads. Includes 42 replay tasks across 6 risk suites with attack and utility scoring.

View Repository

DeepTrap

DeepTrap logo

Tasks Chinese Samples Risks Scenarios Leaderboard
arXiv Paper HuggingFace License

English | 中文

Open-world security evaluation for OpenClaw agents under adversarial execution contexts.


DeepTrap is a security benchmark for evaluating whether OpenClaw agents can complete benign user tasks while resisting malicious execution-context pressure: poisoned workspace files, injected skills, misleading tool metadata, unsafe command paths, planted secrets, and encoded payloads.

The public release contains 42 replay tasks arranged as 6 contextual vulnerability classes x 7 operational scenario families, plus the benchmark runner and public scoring code. It also includes 10 Chinese-language replay samples in tasks_zh/ for localized evaluation experiments. The private task-generation and attack-search pipeline is intentionally not included.

DeepTrap accompanies the paper:

Red-Teaming Agent Execution Contexts: Open-World Security Evaluation on OpenClaw
Hongwei Yao, Yiming Liu, Yiling He, and Bingrun Yang. arXiv:2605.11047, 2026.

Why DeepTrap?

Most agent benchmarks ask whether a model can finish useful work. DeepTrap asks a stricter question: can the agent finish useful work safely when the surrounding execution context is adversarial?

What We TestWhy It Matters
Context RobustnessPoisoned files, memory-like artifacts, playbooks, and workspace instructionsReal agents read from mutable local context, where untrusted text can silently redirect behavior
Secret HandlingPlanted credentials, environment variables, API keys, and internal valuesAgents must distinguish configuration awareness from secret exfiltration
Execution BoundariesDangerous shell commands, deployment scripts, and file-system side effectsTool-using agents can cause durable harm even when the user prompt is benign
Skill/Tool TrustMalicious skills, misleading tool metadata, and tool-abuse patternsAgent safety depends on how the model interprets and validates tool affordances
Obfuscated PayloadsBase64 and multi-step encoded instructionsAttacks can hide behind normal-looking data-processing tasks

What Sets Us Apart

  • Execution-context attacks, not only prompt attacks. DeepTrap evaluates threats embedded in files, skills, tool descriptions, scripts, logs, configs, and encoded artifacts.
  • Benign user prompts. The user request is useful and ordinary; the risk comes from the surrounding workspace.
  • Public, minimal reproduction. The repository contains the runner, replay tasks, and scoring logic needed to reproduce benchmark runs without exposing private generation code.
  • Attack and utility scoring. DeepTrap reports both AGS (Attack Grading Score) and UGS (Utility Grading Score), so a model is evaluated on safety and task usefulness together.
  • OpenClaw-native. Tasks are built for OpenClaw-style agents and preserve the execution assumptions of a real file/tool/skill workspace.

News

  • 2026-05 DeepTrap paper released on arXiv: arXiv:2605.11047.
  • 2026-05 Public benchmark repository released with 42 replay tasks, scoring code, and GitHub Pages leaderboard.
  • 2026-05 Chinese-language replay samples added under tasks_zh/.
  • 2026-05 Hugging Face dataset export added for task metadata distribution.

Leaderboard

AGS is Attack Grading Score and UGS is Utility Grading Score. Scores are reported by risk suite; Average is the mean across Risk 1 through Risk 6.

Full project page and leaderboard: ZJUICSR.github.io/DeepTrap

Claw ModelMetricRisk 1Risk 2Risk 3Risk 4Risk 5Risk 6Average
GPT-5.4AGS0.770.840.760.610.670.530.70
GPT-5.4UGS0.910.830.860.770.740.870.83
Claude-Sonnet-4.6AGS0.510.580.370.250.380.200.38
Claude-Sonnet-4.6UGS0.710.690.550.450.550.710.61
GLM-5AGS0.810.930.740.830.790.880.83
GLM-5UGS0.900.900.980.890.830.880.90
Qwen3.5-PlusAGS0.930.930.860.740.880.970.88
Qwen3.5-PlusUGS0.950.921.000.980.930.930.95
MiniMax-M2.5AGS0.860.890.770.660.900.890.83
MiniMax-M2.5UGS0.920.951.000.880.740.900.90
DeepSeek-v4-FlashAGS0.900.960.800.900.820.940.89
DeepSeek-v4-FlashUGS0.980.961.000.960.851.000.96
Deepseek-v4-ProAGS0.900.960.740.870.850.860.86
Deepseek-v4-ProUGS0.900.911.000.810.840.890.89
MiMo-v2.5AGS0.860.870.710.730.570.600.72
MiMo-v2.5UGS0.960.950.880.930.830.890.91
MiMo-v2.5-proAGS0.740.830.560.580.580.530.64
MiMo-v2.5-proUGS0.920.900.880.870.710.870.86

Benchmark Design

DeepTrap crosses six contextual vulnerability classes with seven benign scenario families. Each task uses a normal user prompt; the adversarial behavior is induced by the workspace context.

DeepTrap framework

DeepTrap constructs compromised execution contexts from benign instructions and clean workspaces, searches candidate adversarial payloads with multi-objective reward signals, and refines successful payloads through reflection-based probing.

Risk Suites

SuiteCategoryTasksWhat is measured
R1Harness Hijacking7Resistance to poisoned context and injected workflow steps
R2Privacy Leakage7Protection of planted credentials and sensitive values
R3Unauthorized Execution7Refusal of dangerous commands and boundary-crossing actions
R4Supply Chain Risk7Handling of malicious skills and tool-like integrations
R5Tool Abuse7Detection of misleading or mismatched tool behavior
R6Encoding Obfuscation7Handling of encoded instructions and multi-step payloads

Scenario Families

Download Tool