
Agentic AI security scanner that reasons like an attacker over source code, confirms exploitable flaws with executable PoCs, and drives test-first fixes through hunt, fix, and verify skills.
[!NOTE] A maintained fork of Capital One's VulnHunter (Apache-2.0) — built to run on any agent harness, not just Claude Code. This fork's focus: harness portability, sandboxed (containerized) exploit validation, and measured-impact PoCs. See Why the changes · What this fork changes · The numbers.
From pattern-matching to provability.
VulnHunter is an open-source, agentic AI security tool that applies proactive, attacker-first analysis directly to source code.
Unlike traditional, passive SAST scanners that flag suspicious patterns and often cause false positives, VulnHunter reasons like an adversary. It identifies which defects are actually exploitable, maps prospective attack paths, and proposes targeted, evidence-backed fixes.
Modern software supply chains are deeply interconnected. A single vulnerability in a widely-used open-source component can ripple across thousands of enterprises simultaneously.
VulnHunter was developed internally at Capital One and open-sourced for the community. This fork carries that work forward — same methodology, reworked to run on any agent harness, with sandboxed (containerized) exploit validation and measured-impact PoCs as the roadmap. See What this fork changes.
Dual-use caution VulnHunter performs dual-use cybersecurity work (vulnerability discovery and exploitation). Expect guardrails: most commercially available models apply dual-use cyber safeguards, and aggressive exploitation behavior can trip rate limits or usage flags. VulnHunter's development and testing ran on open-weight, community-provided models — de-risked, abliterated, and uncensored — which are the models likely to matter for organizational use going forward. Audit only code you own or are otherwise authorized to audit.
[!IMPORTANT] Prerequisites & Model Requirements VulnHunter's methodology is built to run on open-weight, community-provided models — the de-risked, abliterated, uncensored ones organizations can actually deploy. A capable reasoning model is required; the strongest model your harness offers gives the best results, but the methodology does not depend on a specific vendor's frontier model. You supply your own model access.
| Capability | Upstream (Capital One) | This fork | Status |
|---|---|---|---|
| Harness portability | Skills invoke Claude Code specifically; installer targets ~/.claude/skills; model gates hardcode Opus; harness pins claude-opus-4-8 | Skills are harness-portable prompt files (any harness with a skills directory + subagents); VULNHUNT_SKILLS_DIR / VULNHUNT_AGENTS_DIR / VULNHUNT_BIN_DIR / VULNHUNT_HOST_CMD / VULNHUNT_MODEL environment contract; model gates rephrased to "your harness's most capable reasoning model" | Shipped |
| No-guess installer | install.sh assumes ~/.claude/skills | Explicit directories, honors GROK_HOME semantics, writes the vh launcher to VULNHUNT_BIN_DIR/~/.local/bin, installs the vulnhunter-run skill + agent definition; Windows .cmd equivalents updated | Shipped |
vulnhunter-run operator skill | — (absent) | Unattended operator: clone → hunt → find-results → write/validate the scan manifest, with explicit stop rules and no improvisation | Shipped |
| Benchmark/judge hardening | Fixed model + basic retry | Model via environment, retry/backoff configuration, analyze_misses pipeline loss-point tracing, per-finding history tracking | Shipped |
| Harness-neutral report language | Claude-specific prose throughout the skills | Harness-neutral tool language (Agent → subagent, Claude CLI → harness session) | Shipped |
| Sandbox-first exploit validation | Exploit tests may be static traces; runtime choice ad hoc | Docker-first runtime provisioning; the runtime recorded per finding; Medium+ severity must execute | In progress |
| Measured-impact PoCs | PoCs are documents; impact asserted | Executable PoC + impact number in the finding (rows exposed, requests amplified, key-hours stranded) | In progress |
VulnHunter's methodology is host-agnostic by nature: it is prompt procedure, not tool binding. The upstream project grew up inside Claude Code — a coherent choice, and the right first home. But the agent-harness landscape has broadened, and a security methodology that installs into only one of them stops being an audit capability and starts being a vendor feature. This fork makes four changes, each with a reason.
Every skill here is a portable prompt file with an explicit environment contract (VULNHUNT_SKILLS_DIR, VULNHUNT_AGENTS_DIR, VULNHUNT_MODEL, VULNHUNT_HOST_CMD), and the model gates now ask for your harness's most capable reasoning model instead of a specific product. Better means: the same methodology installs into whatever harness your team already runs — and becomes comparable across harnesses in benchmark runs, which is how this fork is developed.
The upstream installer copied skills into ~/.claude/skills unconditionally. On a machine running two harnesses — or a harness with a relocated home — that guess installs into the wrong place, silently. The fork's installer asks, or takes environment variables, and fails loudly with the exact instruction when the answer is missing. Better means: safe on multi-harness machines, correct under relocated homes, loud instead of silent when misconfigured.
The original design already demands falsification and exploit tests. What it left open was how hard to work to actually execute them: static trace, mocked test, or a real containerized server. In one six-run benchmark against a single commit, that discretion produced anywhere from 3 to 42 findings — and opposite verdicts on the same sink, one proven against a mock, one closed by a test against a real server. This fork adds a runtime-provisioning procedure (Docker-first, recorded per finding) and a PoC discipline where impact is measured — rows leaked, ×-amplification, key-hours stranded — not narrated. Better means: a finding's validity no longer depends on which model had the instinct to stand up a container. (In progress — the build plan is on the public roadmap; ask in issues or watch the repo's Discussions.)
New in this fork: vulnhunter-run, an unattended operator that clones, hunts, locates results, and writes and validates the scan manifest with explicit stop rules. The benchmark tooling gains model configuration via environment, retry/backoff knobs, and loss-point analysis for missed findings. Better means: the difference between a tool you demo and a tool you schedule.
We benchmark VulnHunter against itself: six full scans of one real production Go service — same commit, five harness/model stacks. The numbers below are from those runs, and they're why this fork exists.