Security Scanner for Agent Skills
A best-effort security scanner for AI Agent Skills that detects prompt injection, data exfiltration, and malicious code patterns. It combines pattern-based detection (YAML + YARA-X), AST and dataflow analysis, an optional LLM-as-a-judge, and a bounded CEL decision layer over typed detector facts.
Important: This scanner provides best-effort detection, not comprehensive or complete coverage. A scan that returns no findings does not guarantee that a skill is free of all threats. See Scope and Limitations below.
Supports OpenAI Codex Skills and Cursor Agent Skills formats following the Agent Skills specification. With --lenient, also scans non-standard formats such as Claude Code .claude/commands/*.md and flat markdown skill repos.
cel-go v0.32.0 runtime to correlate bounded facts after deterministic detection and before optional LLM analysisJoin the Cisco AI Discord to discuss, share feedback, or connect with the team.
Skill Scanner is a detection tool. It identifies known and probable risk patterns, but it does not certify security.
Key limitations:
The final core + CEL development benchmark contains 5,256 malicious and 1,338
benign MaliciousSkillBench packages. Compared with origin/main, the current
scanner raised F1 from 32.92% to 47.73% and recall from 19.88% to 31.43%, while
reducing benign false-positive rate from 3.59% to 1.05%. Precision is 99.16%.
Five CEL-shadow runs were exact and deterministic; CEL evaluated 154 candidates
without proposing a suppression or falling back.
The locked source-disjoint split is weaker: TP=65, FP=42, TN=503, and FN=774,
for 60.75% precision, 7.75% recall, 13.74% F1, and 7.71% FPR. This improves F1
over origin/main (7.40%) but regresses FPR (3.67%), so it does not pass the
promotion gate. Every bundled CEL rule therefore remains in shadow; this
change does not promote any CEL suppression.
Compatibility and supplemental checks found identical CEL-OFF/CEL-SHADOW findings on 111 official Codex, Claude Code, and Cursor skills (30 MEDIUM+ and 8 HIGH/CRITICAL packages), with five stable runs. The NotInject hard-negative set had 0/339 actionable matches. HarmfulSkillBench had 7/200 actionable and 6/200 HIGH+ packages with one quarantined sample, while OpenSkillRisk had 76/263 actionable packages with two host quarantines. The latter two are positive-only recall diagnostics and cannot measure precision or FPR. Optional ATR results are outside this release scope. See Detection Evaluation and Rollout for methodology, provenance, confidence intervals, and limitations.