Back to updates
New releaseAug 4, 2026

skill-scanner v2.0.13

Multi-engine security scanner for AI agent skills detecting prompt injection, data exfiltration, and malicious code via static analysis, behavioral dataflow, and LLM semantic analysis with CI/CD integration.

Share

Skill Scanner

License CPython 3.11–3.14 PyPI version CI Discord Cisco AI Defense AI Security Framework Ask DeepWiki

A best-effort security scanner for AI Agent Skills that detects prompt injection, data exfiltration, and malicious code patterns. It combines pattern-based detection (YAML + YARA-X), AST and dataflow analysis, an optional LLM-as-a-judge, and a bounded CEL decision layer over typed detector facts.

Important: This scanner provides best-effort detection, not comprehensive or complete coverage. A scan that returns no findings does not guarantee that a skill is free of all threats. See Scope and Limitations below.

Supports OpenAI Codex Skills and Cursor Agent Skills formats following the Agent Skills specification. With --lenient, also scans non-standard formats such as Claude Code .claude/commands/*.md and flat markdown skill repos.


Highlights

  • Multi-Engine Detection - Static analysis, behavioral dataflow, LLM semantic analysis, and cloud-based scanning for layered, best-effort coverage
  • Typed CEL Decisions - The core scanner uses the official cel-go v0.32.0 runtime to correlate bounded facts after deterministic detection and before optional LLM analysis
  • Finding Review - The optional Meta-analyzer correlates, prioritizes, and can filter findings; paired accuracy validation remains pending
  • CI/CD Ready - SARIF output for GitHub Code Scanning, reusable GitHub Actions workflow, exit codes for build failures
  • Pre-commit Hook - Standard pre-commit framework integration to scan skills before every commit
  • Extensible - Plugin architecture for custom analyzers

Join the Cisco AI Discord to discuss, share feedback, or connect with the team.


Scope and Limitations

Skill Scanner is a detection tool. It identifies known and probable risk patterns, but it does not certify security.

Key limitations:

  • No findings ≠ no risk. A scan that returns "No findings" indicates that no known threat patterns were detected. It does not guarantee that a skill is secure, benign, or free of vulnerabilities.
  • Coverage is inherently incomplete. The scanner combines signature-based detection, LLM-based semantic analysis, behavioral dataflow analysis, optional cloud services, and configurable rule packs. While this approach improves coverage, no automated tool can detect every technique, especially novel or zero-day attacks.
  • False positives and false negatives can occur. Consensus modes and meta-analysis can help review findings, but no configuration eliminates all incorrect classifications. Tune the scan policy to your risk tolerance.
  • Human review remains essential. Automated scanning is one component of a defense-in-depth strategy. High-risk or production deployments should pair scanner results with manual code review and/or threat modeling.

Current modernization evidence

The final core + CEL development benchmark contains 5,256 malicious and 1,338 benign MaliciousSkillBench packages. Compared with origin/main, the current scanner raised F1 from 32.92% to 47.73% and recall from 19.88% to 31.43%, while reducing benign false-positive rate from 3.59% to 1.05%. Precision is 99.16%. Five CEL-shadow runs were exact and deterministic; CEL evaluated 154 candidates without proposing a suppression or falling back.

The locked source-disjoint split is weaker: TP=65, FP=42, TN=503, and FN=774, for 60.75% precision, 7.75% recall, 13.74% F1, and 7.71% FPR. This improves F1 over origin/main (7.40%) but regresses FPR (3.67%), so it does not pass the promotion gate. Every bundled CEL rule therefore remains in shadow; this change does not promote any CEL suppression.

Compatibility and supplemental checks found identical CEL-OFF/CEL-SHADOW findings on 111 official Codex, Claude Code, and Cursor skills (30 MEDIUM+ and 8 HIGH/CRITICAL packages), with five stable runs. The NotInject hard-negative set had 0/339 actionable matches. HarmfulSkillBench had 7/200 actionable and 6/200 HIGH+ packages with one quarantined sample, while OpenSkillRisk had 76/263 actionable packages with two host quarantines. The latter two are positive-only recall diagnostics and cannot measure precision or FPR. Optional ATR results are outside this release scope. See Detection Evaluation and Rollout for methodology, provenance, confidence intervals, and limitations.


Documentation

Categories