Multi-agent static application-security review harness for AI coding agents: maps codebases, hunts vulnerability classes, chains and verifies findings, and reports to SARIF, JSON, and PDF.
A multi-agent application-security review harness for Claude Code (and, later, other AI agents). One router skill dispatches to a full offensive-security pipeline that maps a codebase, hunts vulnerabilities with a per-class knowledge base, chains findings into escalations, verifies real impact, and reports to README / JSON / SARIF / doc / PDF.
Scope: this harness performs static analysis (source review, data-flow tracing, PoC/payload construction) on code you own or are authorized to test. It does not attack live third-party systems.
security-harness/ # a plugin marketplace
└── plugins/security-harness/
├── skills/
│ ├── sh-router # single entry point - routes any appsec request
│ ├── sh-security-review # the pipeline orchestrator (Stages 0-5)
│ └── sh-kb-* (15) # per-vuln-class knowledge bases
├── agents/
│ ├── sh-recon # map: Graft graph + stack/SBOM/CVE + attack surface
│ ├── sh-hunter # find: source→sink hunting, one per class (parallel)
│ ├── sh-chainer # escalate: combine findings into attack chains
│ ├── sh-verifier # confirm: offensive + seceng + dev verification + PoC
│ └── sh-reporter # deliver: README/JSON/SARIF/HTML/PDF/doc
└── references/ # shared contracts (finding schema, SARIF map, state files, rubrics)
sh-kb-* skills)access-control (IDOR/BOLA/priv-esc) · sqli · xss · ssrf · injection (cmd/code/SSTI/LDAP) · auth (session/JWT) · deserialization · path-traversal (LFI/RFI) · secrets · csrf · xxe · open-redirect · crypto · race-conditions · file-upload. Dependency CVEs/SBOM are handled by the recon stage.
The harness ships for several agents. Full packaging details and the release
checklist are in docs/DISTRIBUTION.md.
Claude Code - add this repo as a plugin marketplace and install the plugin:
/plugin marketplace add dmdhrumilmistry/security-harness
/plugin install security-harness
/plugin marketplace add accepts any of: a GitHub owner/repo (as above), a full git URL
(https://github.com/dmdhrumilmistry/security-harness.git), or a local path to a clone
(e.g. /plugin marketplace add ./security-harness from the directory containing your checkout).
Then run /plugin install security-harness and reload when prompted.
Gemini CLI - a native extension, manifest at the repo root:
gemini extensions install https://github.com/dmdhrumilmistry/security-harness
opencode, Codex, or any agentskills.io agent - copy the
skills into a discovery directory. Codex additionally picks up AGENTS.md on its own:
git clone https://github.com/dmdhrumilmistry/security-harness
cd security-harness
python3 scripts/sync-agent-skills.py --install agents # ~/.agents/skills
python3 scripts/sync-agent-skills.py --install opencode # ~/.config/opencode/skills
Graft is installed and set up automatically by the pipeline. Stage 0 runs npm install -g @nanonets/graft if it's missing (requires Node/npm), then graft init <target> --no-agents --no-global
to register the Graft MCP server and freshness hooks for the target repo. The graph itself (<target>/graft/,
auto-gitignored) is built during recon. To pre-install manually: npm install -g @nanonets/graft. Graft's
structural build is free and needs no API key; the optional --deep LLM pass uses GRAFT_API_KEY /
GRAFT_PROVIDER / GRAFT_MODEL when set.
The other tools are also auto-installed by Stage 0 when missing (via whatever package manager is on the
machine - winget/choco/scoop, brew, apt, npm/pip/go - see references/tooling-setup.md). Installs are
announced, prefer no-elevation methods, and never block the run: anything that can't be installed is simply
marked unavailable and the pipeline falls back. Stage 0 installs only what fills a missing capability group:
syft · CVEs: one of grype
(preferred), trivy, or osv-scannerwkhtmltopdf or pandoc (for PDF/DOCX); otherwise you get report.html (or a headless-Chrome PDF).All of them are optional - the pipeline degrades gracefully to native search + manifest parsing if none install.
Invoke the router with a natural-language request:
/sh-router full security review of ./api
/sh-router find SQLi and IDOR in src/
/sh-router just map this codebase # recon only
Or call the pipeline directly:
/sh-security-review . classes:sqli,access-control,ssrf depth:deep
/sh-security-review . stage:report # regenerate reports for the latest run
Each stage runs on a model matched to its cognitive load, so tokens are spent where discovery quality actually depends on them and saved on mechanical work. This is the default - no arguments needed.
| Stage | Default model |
|---|---|
| recon | sonnet |
| hunt (per class) | haiku for pattern classes (secrets, crypto, open-redirect, csrf) · sonnet for source→sink tracing (sqli, xss, ssrf, injection, path-traversal, xxe, file-upload, auth) · opus for deep-logic classes (access-control, race-conditions, deserialization) |
| chain | opus |
| verify | opus (the precision gate - kept strong) |
| report | haiku |
Override with the models: argument (passed through the router too):
/sh-security-review . # default tiered map above
/sh-security-review . models:max # every stage + hunter on opus (max quality, max cost)
/sh-security-review . models:cheap # aggressive downshift (trades some verify precision)
/sh-security-review . models:verify=opus,hunt=sonnet # per-stage overrides
/sh-security-review . models:report=sonnet,hunt.pattern=sonnet # per-hunter-tier override
Stages: setup, recon, hunt, chain, verify, report. Models: opus, sonnet, haiku, inherit. For hunt,
a bare model flattens all hunters to it; hunt.pattern / hunt.trace / hunt.logic target one tier.
Other token savers are built in: recon only spawns hunters for classes with real attack surface, hunters
query the Graft graph instead of reading whole files, and findings.json/SARIF are generated by a
deterministic script rather than the model.
Everything lands under <target>/.security-harness/<run-id>/:
recon.md, codebase-map.json - the map (stack, SBOM, CVEs, attack surface).findings.jsonl → chains.md → verified.jsonl - the working state (see references/state-files.md).reports/ - README.md, findings.json, results.sarif, report.html, report.pdf (+ report.docx).Each published finding carries a payload, a PoC, the verification verdict, CWE/OWASP ids, CVSS, and a code-level mitigation.
sh-pr-review reviews a single pull request rather than a whole codebase, posts the
result as inline comments on the PR itself, and sets a security/pr-review commit
status that branch protection can enforce.
Run it from your own machine, on any PR you can read. Install the plugin and ask:
review https://github.com/acme/api/pull/128
review PR 42
security review this PR
Paste a PR link and it reviews that PR in that repository, cloning it to a temporary directory first, because the hunters read files and not just the patch. Nothing is written into the repo you are working in.
Pass a bare number and it resolves against the repo you are currently in, the one
git remote points at. Pass nothing and it takes the open PR for your current branch.
Phase 7 prints the findings and the verdict and asks before posting anything - a decline is a normal outcome, and the payload stays on disk for you to post later.
Before spending any analysis it checks whether you can actually write to the target repo, so reviewing someone else's project tells you up front that posting will 403 rather than discovering it ten minutes in.
Three properties make it usable as a merge gate rather than noise:
pr_impact
of introduced, aggravated, or pre_existing. The first two block; pre_existing
is reported and never blocks. Blocking a merge over code the author never wrote is how
a required check gets deleted, so when a hunter is unsure between aggravated and
pre_existing, it must pick pre_existing.sh-kb-*
bases use, then picks a tier. Tier 0 (no security-relevant change) launches nothing at
all and still sets the status. Tier 3 runs the full pipeline.| Verdict | Status | When |
|---|---|---|
| fail | failure | introduced or aggravated finding at or above --fail-on (default medium), confidence >= 80 |
| warn | success | nothing introduced or aggravated; pre-existing findings reported |
| pass | success | no findings, or triage stopped at Tier 0 |
| error | error | the review could not complete |
warn reports success on purpose: a warning that blocks a merge is a failure with
extra steps, and teams respond by removing the check. error is kept distinct from
failure so a broken run never looks like a vulnerability it did not find.
The review event is always COMMENT, never REQUEST_CHANGES or APPROVE. The commit
status is the enforcement mechanism, and it is the one branch protection reads.
Scope: the skill writes to the pull request and the commit status, and nowhere else. It opens no issues and creates nothing in any external tracker.
A PR gets reviewed once per push, so the second review has to be cheaper than the first or the tool becomes something people turn off.
Dedup happens before the spending, not before the posting. The fingerprints already on the PR are read in Phase 1 and handed to the hunters and the verifier. Finding a duplicate at the end would mean the most expensive model in the pipeline had already re-confirmed a conclusion that was written on the PR the whole time. This needs no cache: the state lives in the PR, so it works on a cold machine and in CI.
A local cache makes the rest incremental. sh-review-cache stores each run's file
hashes, findings and verdicts under your OS cache directory (never in the repo, since a
cross-repo review runs in a temp clone that gets deleted). The next review re-hunts only
files whose content actually changed, reuses verdicts for findings that are unchanged, and
reuses the recon map if nothing it covers moved.
A base-branch merge costs nothing. Merging main into a PR branch changes the head
SHA and nothing the author wrote, but a commit status is pinned to a SHA, so the required
check silently disappears from the new head. When the PR's own files are byte-identical
and the base delta touches nothing the findings depend on, the previous verdict is
re-stamped onto the new SHA with no agents launched at all. That last condition is what
makes it safe: a base merge that deletes a sanitizer leaves every PR file unchanged while
turning a safe line into an exploitable one.
Invalidation is deliberately conservative, because a stale entry in a security tool does
not make it slow, it makes it wrong. The cache key hashes every sh-kb-* knowledge
base, so a KB update invalidates every cached finding - a cached "clean" must never
suppress the finding that update was written to catch. Model identity, skill version, file
content and a 7-day TTL all invalidate too, and an unspecified model is treated as a miss.
--no-cache disables it, --refresh-cache re-baselines, and run.md records per phase
what was launched, reused and skipped, so a cache that quietly stops hitting is visible
rather than assumed.
The review and the commit status are posted by default. A review that was computed and
never delivered helped nobody. --confirm restores a prompt before posting, --dry-run
sends nothing, --no-status posts the review but leaves the commit status alone.
Posting goes through scripts/sh-pr-post.py rather than hand-built API calls, because it
is a multi-step operation with a mandatory tail: review, then status, then receipts, with
a 422 recovered by moving the comment rather than shifting a line number. The script
never exits leaving the status at pending - if the review cannot be posted it still
sets error, saying the tooling failed rather than accusing the PR.
Inline comments are reserved for findings at medium or above with confidence >= 80. Low-severity findings go in the collapsed body section, so a low-priority finding showing up with zero inline comments is the policy working, not a failure.
Every run records what it cost, so "the cache is working" and "reviews got slower" stop being matters of opinion.
python3 <skill>/scripts/sh-metrics.py path # where records live
python3 <skill>/scripts/sh-metrics.py report # aggregate, by model
python3 <skill>/scripts/sh-metrics.py purge --older-than-days 30
Two append-only JSONL files - runs.jsonl (repo, PR, tier, verdict, totals, the flags you
passed) and events.jsonl (one line per phase or agent: model, tokens, duration, outcome,
whether it was reused from cache). JSONL so a crashed run still leaves valid lines above
the crash.
| Platform | Metrics | Cache |
|---|---|---|
| Linux / BSD | $XDG_DATA_HOME/security-harness/metricsdefault ~/.local/share/security-harness/metrics | $XDG_CACHE_HOME/security-harnessdefault ~/.cache/security-harness |
| macOS | ~/Library/Application Support/security-harness/metrics | ~/Library/Caches/security-harness |
| Windows | %LOCALAPPDATA%\security-harness\metrics | %LOCALAPPDATA%\security-harness\cache |
Linux follows the XDG Base Directory spec, so both honour XDG_DATA_HOME and
XDG_CACHE_HOME when set and fall back to ~/.local/share and ~/.cache when they are
not. Override either directly with SH_METRICS_DIR and SH_REVIEW_CACHE_DIR.
On python vs python3: most Linux distributions ship python3 and have no python
at all, so the examples here use python3. The bundled scripts carry a
#!/usr/bin/env python3 shebang and are executable, so ./scripts/sh-metrics.py report
works directly on Linux and macOS. The skill resolves
PY="$(command -v python3 || command -v python)" once per run, which covers all three
platforms including Git Bash on Windows.
Strictly local. Neither script contains any network code or reporting endpoint. Anything token-shaped is redacted before it is written, because local files get pasted into issues.
Optional, and a separate decision from using the skill. Run it by hand on your own PRs for a while first, so you know what it says about your codebase before it says it in front of your team.
When you are ready, "Enforcing the check on a repository" in
references/pr-review-mapping.md
has a copy-paste workflow for your repo, plus the fail-safe that stops a dead job
leaving a required check stuck on pending.
The threshold stays at the default medium. Pre-existing findings never block a merge,
so an unscanned codebase does not produce a wall of red on day one - only what a PR
actually introduces or aggravates can fail it.
sh-recon) - build the Graft graph; detect stack/versions; SBOM + CVEs; enumerate entry
points, trust boundaries, and dangerous sinks.sh-hunter ×N, parallel) - one hunter per relevant class loads its sh-kb-* knowledge base,
traces attacker input from source to sink, and records candidates. A shared attempts ledger stops
agents from repeating each other's probes.sh-chainer) - compose findings into higher-severity attack paths.sh-verifier) - refute first, then confirm exploitability from evidence, build PoCs, assign
CVSS, and cut false positives.sh-reporter) - produce the deliverables.Subagents share nothing but files; the contract is in plugins/security-harness/references/state-files.md.
Add a new vulnerability class by creating skills/sh-kb-<class>/SKILL.md following the shared template
(When to hunt · Sources & sinks · Detection recipe · Payloads/PoC · False-positive filters · CWE/OWASP ·
Chaining hints · Mitigation), then add its slug to the class enum in references/finding-schema.json
and the routing table in skills/sh-router/SKILL.md.
A scheduled GitHub Action (.github/workflows/update-knowledge-base.yml) keeps the sh-kb-* knowledge
bases fresh. Every alternate day (and on manual workflow_dispatch), it runs an agent to distill new,
reputable public security research - OWASP, PortSwigger Research, CWE/CAPEC, NIST, MDN, curated GitHub
repos, and public HackerOne disclosures - into small, well-sourced improvements. A second, adversarial
reviewer agent then scans the resulting diff for malicious/injected content, and the PR is
auto-merged only if that reviewer approves.
PR creation is deliberately separated from review and merge, so the thing that writes the diff is never the thing that decides to ship it.
Stage 1 - update-knowledge-base.yml
(scheduled or manual). One job, create-pr:
automated/kb-update
branch, labeled awaiting-review.Stage 2 - kb-review-and-merge.yml
(dispatched by stage 1, or run by hand against any automated PR). Two jobs:
review - a separate agent run inspects the diff adversarially for
prompt-injection artifacts, out-of-scope edits, secrets/exfil, PII, weaponized
exploits, off-allowlist sourcing, or house-style violations. It has no web and no
shell, and fails closed: anything suspicious, any uncertainty, or a missing
verdict file → REJECT. The verdict is posted as a PR comment and drives the label.merge - runs only on APPROVE, and merges the PR. A REJECT skips it and
the blocked job reports why.Why a dispatch rather than a
pull_requesttrigger: a PR opened byGITHUB_TOKENdoes not triggerpull_requestworkflows.workflow_dispatchis one of the two events exempt from that recursion guard, so stage 1 can hand off reliably.
Auto-merge means an approving agent lands code in main. The controls on that:
automated/* (ALLOWED_HEAD_PREFIX in the workflow).main requiring an approving review, the PR queues and waits for a human instead of
merging. It falls back to an immediate merge only on repos where auto-merge is off.auto_merge input to false on a manual run to review without merging.Requires the repo setting "Allow GitHub Actions to create and approve pull requests" (Settings → Actions → General → Workflow permissions) so the workflow can open the PR. If you want a human in the loop despite auto-merge, protect
mainwith a branch-protection rule requiring a pull request and at least one approving review - the auto-merge path honours it.
Both stages run through .github/actions/ai-agent,
a composite action that dispatches to whichever agent you configure. Claude Code, OpenAI
Codex, Gemini CLI, and an escape hatch for anything else:
agent | Runs | Credential |
|---|---|---|
claude (default) | anthropics/claude-code-action@v1 | CLAUDE_CODE_OAUTH_TOKEN or ANTHROPIC_API_KEY |
codex | codex exec --full-auto | OPENAI_API_KEY |
gemini | gemini --yolo --prompt | GEMINI_API_KEY |
custom | your KB_AGENT_INSTALL / KB_AGENT_COMMAND | whatever it needs |
Pick per run from the workflow_dispatch inputs, or set repo variables to change the
default: KB_AGENT and KB_MODEL for the generator, KB_REVIEW_AGENT and
KB_REVIEW_MODEL for the reviewer. Running the generator and the reviewer on different
agents is a meaningful hardening step: an injection tuned for one model is less likely
to land on a second, independent one.
For agent: custom, set KB_AGENT_COMMAND to a shell command. The prompt is written to
the file named by $AGENT_PROMPT_FILE, and $AGENT_MODEL carries the model input.
Prompt-injection defenses, since the generator reads the open web:
WebFetch is restricted to the trusted domains in
.github/kb-update/trusted-sources.md (mirrored in the workflow's --allowedTools). WebSearch can
discover URLs, but only allowlisted domains can actually be fetched..github/kb-update/prompt.md) instructs Claude to
treat every fetched byte as untrusted reference material and to ignore any instructions embedded in a
page - HackerOne report bodies (user-generated) are flagged as the highest-risk tier..github/kb-update/review-prompt.md) is what stands between fetched content
and main - nothing merges without its explicit approval.KB_AGENT and KB_REVIEW_AGENT to two different engines.Setup:
CLAUDE_CODE_OAUTH_TOKEN (default), ANTHROPIC_API_KEY, OPENAI_API_KEY, or GEMINI_API_KEY.
The OAuth token authenticates against your Claude subscription's usage limits rather than a metered
API key - generate it locally with claude setup-token (requires an active Claude Pro/Max subscription)
and paste the result.main
requiring a PR and an approving review, so no automated change can land without a human even with
auto-merge on.trusted-sources.md and the matching
WebFetch(domain:...) entries in the workflow - keep the two in sync.Each run records what it did in .github/kb-update/last-run-summary.md.
AGENTS.md, a Gemini CLI extension, and scripts/sync-agent-skills.py
for .agents/skills and opencode. See docs/DISTRIBUTION.md.needs-runtime findings.MIT