
skill-scanner v2.0.13
Multi-engine security scanner for AI agent skills detecting prompt injection, data exfiltration, and malicious code via static analysis, behavioral dataflow, and LLM semantic analysis with CI/CD integration.
Skill Scanner
A best-effort security scanner for AI Agent Skills that detects prompt injection, data exfiltration, and malicious code patterns. It combines pattern-based detection (YAML + YARA-X), AST and dataflow analysis, an optional LLM-as-a-judge, and a bounded CEL decision layer over typed detector facts.
Important: This scanner provides best-effort detection, not comprehensive or complete coverage. A scan that returns no findings does not guarantee that a skill is free of all threats. See Scope and Limitations below.
Supports OpenAI Codex Skills and Cursor Agent Skills formats following the Agent Skills specification. With --lenient, also scans non-standard formats such as Claude Code .claude/commands/*.md and flat markdown skill repos.
Highlights
- Multi-Engine Detection - Static analysis, behavioral dataflow, LLM semantic analysis, and cloud-based scanning for layered, best-effort coverage
- Typed CEL Decisions - The core scanner uses the official
cel-gov0.32.0 runtime to correlate bounded facts after deterministic detection and before optional LLM analysis - Finding Review - The optional Meta-analyzer correlates, prioritizes, and can filter findings; paired accuracy validation remains pending
- CI/CD Ready - SARIF output for GitHub Code Scanning, reusable GitHub Actions workflow, exit codes for build failures
- Pre-commit Hook - Standard pre-commit framework integration to scan skills before every commit
- Extensible - Plugin architecture for custom analyzers
Join the Cisco AI Discord to discuss, share feedback, or connect with the team.
Scope and Limitations
Skill Scanner is a detection tool. It identifies known and probable risk patterns, but it does not certify security.
Key limitations:
- No findings ≠ no risk. A scan that returns "No findings" indicates that no known threat patterns were detected. It does not guarantee that a skill is secure, benign, or free of vulnerabilities.
- Coverage is inherently incomplete. The scanner combines signature-based detection, LLM-based semantic analysis, behavioral dataflow analysis, optional cloud services, and configurable rule packs. While this approach improves coverage, no automated tool can detect every technique, especially novel or zero-day attacks.
- False positives and false negatives can occur. Consensus modes and meta-analysis can help review findings, but no configuration eliminates all incorrect classifications. Tune the scan policy to your risk tolerance.
- Human review remains essential. Automated scanning is one component of a defense-in-depth strategy. High-risk or production deployments should pair scanner results with manual code review and/or threat modeling.
Current modernization evidence
The final core + CEL development benchmark contains 5,256 malicious and 1,338
benign MaliciousSkillBench packages. Compared with origin/main, the current
scanner raised F1 from 32.92% to 47.73% and recall from 19.88% to 31.43%, while
reducing benign false-positive rate from 3.59% to 1.05%. Precision is 99.16%.
Five CEL-shadow runs were exact and deterministic; CEL evaluated 154 candidates
without proposing a suppression or falling back.
The locked source-disjoint split is weaker: TP=65, FP=42, TN=503, and FN=774,
for 60.75% precision, 7.75% recall, 13.74% F1, and 7.71% FPR. This improves F1
over origin/main (7.40%) but regresses FPR (3.67%), so it does not pass the
promotion gate. Every bundled CEL rule therefore remains in shadow; this
change does not promote any CEL suppression.
Compatibility and supplemental checks found identical CEL-OFF/CEL-SHADOW findings on 111 official Codex, Claude Code, and Cursor skills (30 MEDIUM+ and 8 HIGH/CRITICAL packages), with five stable runs. The NotInject hard-negative set had 0/339 actionable matches. HarmfulSkillBench had 7/200 actionable and 6/200 HIGH+ packages with one quarantined sample, while OpenSkillRisk had 76/263 actionable packages with two host quarantines. The latter two are positive-only recall diagnostics and cannot measure precision or FPR. Optional ATR results are outside this release scope. See Detection Evaluation and Rollout for methodology, provenance, confidence intervals, and limitations.