
AI-driven pentest harness with black-box, white-box, grey-box, host/cloud, and LLM red-team modes; validates findings with cross-model voting and tool receipts.
Autonomous, multi-model penetration-testing harness — Rust, CLI-only.
by Joas A Santos & Red Team Leaders
⭐ If this is useful, star the repo — it helps a lot.
📖 New here? Read the full Tutorial & User Guide → — every mode, flag, config and example explained. Version-by-version changes live in RELEASE.md.
NeuroSploit turns a URL, a source repository, a running app, or a host/IP into
an autonomous security engagement. A Rust harness (tokio) drives a pool of
LLMs — via API key or local subscription (Claude Code / Codex / Gemini /
Grok) — recons the target, intelligently selects only the agents that match the
discovered surface, runs them in parallel, chains findings into deeper
impact, and validates every claim by cross-model voting + tool-receipt
grounding before reporting. It ships 435 markdown agents and a Mission
Control TUI.
may_assert
gate is a mathematical anti-hallucination rule (don't claim exploitability
while the belief is diffuse).file:line into the reviewed source — a code citation is the receipt) for
white-box SAST & skills audits, and either for grey-box; ungrounded claims
are demoted.aws/gcloud/az). Connect via creds.yaml: AWS keys, a Google
service-account JSON, or an Azure service principal — see
.This is the slim, Rust-only distribution (
neurosploit-rs/+agents_md/). The earlier Python engine and web GUIs live on the olderv3.4.0branch.
Linux / macOS (x64 & arm64):
curl -fsSL https://raw.githubusercontent.com/JoasASantos/NeuroSploit/main/setup.sh | bash
Windows (PowerShell, x64 & arm64):
irm https://raw.githubusercontent.com/JoasASantos/NeuroSploit/main/install.ps1 | iex
| OS | x64 | arm64 |
|---|---|---|
| Linux (Kali recommended) | ✅ | ✅ |
Pure Rust + stdlib, so it builds natively everywhere a stable Rust toolchain runs.
The installer auto-detects OS/arch and installs Rust if missing. On native Windows
use install.ps1; under WSL2 / Git Bash the setup.sh one-liner also works.
The installer auto-installs Rust if needed, clones the repo to ~/.neurosploit,
builds the release binary, and links neurosploit into ~/.local/bin. Re-run it
any time to update. Tweak with env vars: NEUROSPLOIT_REF (branch/tag),
NEUROSPLOIT_DIR, PREFIX.
Prefer to build by hand?
git clone https://github.com/JoasASantos/NeuroSploit && cd NeuroSploit/neurosploit-rs
cargo build --release # → target/release/neurosploit
# easiest path — just run it; the interactive session asks everything:
neurosploit
# or one-liner (subscription login, no API key needed):
neurosploit run http://testphp.vulnweb.com/ --subscription --model anthropic:claude-opus-4-8 -v
# white-box — review a source repository (SAST agents, file:line evidence):
git clone https://github.com/digininja/DVWA /tmp/DVWA
neurosploit whitebox /tmp/DVWA --subscription --model anthropic:claude-opus-4-8 -v
# grey-box — review the code AND exploit the running app together:
neurosploit greybox /tmp/DVWA --url http://localhost:8080/ --creds creds.yaml \
--subscription --model anthropic:claude-opus-4-8 --mcp -v
# host / infra — Linux / Windows / Active Directory (SSH/Win creds in creds.yaml):
neurosploit host 10.0.0.10 --creds creds.yaml --subscription --model anthropic:claude-opus-4-8 -v
# 🛰 Mission Control TUI — live panels (header/feed/findings/targets) + a composer
# you can type in WHILE the run streams (summary · pause · errors · notes):
neurosploit tui http://testphp.vulnweb.com/ --subscription --model anthropic:claude-opus-4-8 --mcp
Full step-by-step for every mode (black/white/grey/host) is in TUTORIAL.md.
No login? Use an API key instead — see Authentication.
A browser UI for the same harness — every action spawns the real compiled CLI and parses its output; nothing about the harness logic is reimplemented in the browser.
cd neurosploit-rs && cargo build --release # once
node web/server.js # → http://localhost:4173
Zero npm dependencies (Node built-ins only).
claude CLI (Opus, your Anthropic subscription) to generate an actual specialist-agent
markdown file into agents_md/vulns/, in the same format every built-in agent uses, pinnable
immediately. Falls back to a plain focus-text hint if Claude isn't available.chains_from when the harness
set one) instead of a flat list — click any node or row for the full finding detail, including
any PoC script the exploiting agent wrote to pocs/.creds.yaml for the run), and per-provider API keys held in the
server process's memory only — never written to disk.Full API reference: web/API.md · quick start: web/README.md.
Wire NeuroSploit into your SDLC. Toggle from the REPL (/integrations) or the CLI
(neurosploit integrations enable github|gitlab|jira). Tokens are never stored
— only the name of the env var is saved; the value is read from your environment.
export GITHUB_TOKEN=ghp_... # PAT with `repo` scope (private repos)
neurosploit integrations enable github
# Review a Pull Request's code (clones the PR head, white-box) and comment back:
neurosploit pr digininja/DVWA 42 --subscription --model anthropic:claude-opus-4-8 --comment
# Same, but BLOCK the merge on a confirmed critical: fails the check, sets a
# `neurosploit/security` commit status, and posts a REQUEST_CHANGES review.
neurosploit pr digininja/DVWA 42 --model anthropic:claude-opus-4-8 --comment --fail-on critical
# Watch a branch and re-review on every new commit:
neurosploit watch myorg/private-app --branch main --subscription --model anthropic:claude-opus-4-8
# Private GitLab repo (token-injected clone) — works in whitebox/greybox:
export GITLAB_TOKEN=glpat-... ; neurosploit integrations enable gitlab
neurosploit whitebox https://gitlab.com/myorg/private-svc --subscription --model anthropic:claude-opus-4-8
# Open a Jira card per finding (any engagement):
export [email protected] JIRA_API_TOKEN=... # set base/project once: /integrations setup jira
neurosploit whitebox https://github.com/myorg/app --jira --subscription --model anthropic:claude-opus-4-8
Two ready-made workflows ship in examples/github-actions/ — copy
them into your repo:
neurosploit-pr-gate.yml — reviews every PR and blocks the merge on a
confirmed critical. Make it enforcing: Settings → Branches → require the
neurosploit-pr-gate status check (and/or require review to honor the
REQUEST_CHANGES). Set ANTHROPIC_API_KEY (or swap the model) in Actions secrets;
the built-in GITHUB_TOKEN covers statuses/reviews.neurosploit-mention.yml — comment @neurosploit on a PR or issue to
trigger a scan (only repo writers can). Text after the mention is the
instruction (any language): @neurosploit focus SQLi and IDOR, or
@neurosploit scan https://staging.app for a black-box run.📖 Step-by-step setup for each tool: TUTORIAL-INTEGRATION.md.
Add a cloud block to creds.yaml and the harness exports the right env vars so
the AWS/GCP/Azure agents can drive aws / gcloud / az. Secrets stay in your
file/secret-manager; agents do read-only enumeration first, never destructive.
# --- AWS: static keys (or a named profile) ---
aws:
access_key_id: AKIA...
secret_access_key: ...
# session_token: ... # if using temporary creds
region: us-east-1
# profile: my-sso-profile # alternative to keys
# --- GCP: service-account JSON (path recommended; inline single-line also works) ---
gcp:
service_account_json: /path/to/sa.json
project: my-project-id
# --- Azure: service principal (recommended for automation) ---
azure:
tenant_id: ...
client_id: ...
client_secret: ...
subscription_id: ...
neurosploit host my-cloud-account --creds creds.yaml \
--subscription --model anthropic:claude-opus-4-8 -v
Agents cover IAM privilege-escalation, storage exposure (S3/GCS/Blob), compute &
network exposure, secrets (Secrets Manager / Secret Manager / Key Vault),
service-account/SP abuse, and identity enumeration (Entra ID). Best-practice
auth: AWS access keys or profile; GCP a service-account JSON
(GOOGLE_APPLICATION_CREDENTIALS); Azure a service principal
(az login --service-principal).
Give NeuroSploit two or more named roles in creds.yaml and it authenticates
as each and tests cross-role access (a low-priv role reaching another user's
object or an admin function is a finding):
admin:
jwt: eyJ... # per role: jwt | header (raw) | cookie | apikey | login+username+password
user:
apikey: abc123 # → X-Api-Key: abc123
victim:
cookie: "session=deadbeef"
neurosploit run https://app.example --creds creds.yaml \
--subscription --model anthropic:claude-opus-4-8 -v
Each finding is proven with the authorized vs unauthorized request pair, under the data-safety guardrail (read-only, PII masked).
Every request is tagged with an identifying User-Agent (default
NeuroSploit/<ver> …, change with /ua or NEUROSPLOIT_UA) plus an
X-NeuroSploit-Scan header, and every finding is stamped "Identified and
validated by NeuroSploit" — so provenance travels in the traffic, the finding
text, findings.json and the report footer.
cd neurosploit-rs
cargo build --release # → target/release/neurosploit
Requires a Rust toolchain (rustup). Recommended: run on Kali Linux (or the
Kali Docker image) so the offensive tools the agents use are already present:
docker run -it --rm kalilinux/kali-rolling
apt update && apt install -y curl nmap ffuf nodejs npm
# rustscan (faster port scan): cargo install rustscan (or grab a release from GitHub)
The agents degrade gracefully: if rustscan isn't installed they use nmap; if
neither, they probe with curl. If a Playwright MCP browser is available they use
it for JS-heavy pages, otherwise they fall back to curl.
Run with no arguments for an interactive wizard:
./target/release/neurosploit
Or drive it directly:
# Black-box — subscription (no API key), Opus, browser via Playwright if present, verbose
./target/release/neurosploit run http://testphp.vulnweb.com/ \
--subscription --model anthropic:claude-opus-4-8 --mcp -v
# Black-box — API keys, multi-model voting panel (1st finds, others adjudicate)
./target/release/neurosploit run http://testphp.vulnweb.com/ \
--model anthropic:claude-opus-4-8 --model openai:gpt-5.1 --vote-n 3
# White-box — clone a vulnerable app and review its source
git clone https://github.com/digininja/DVWA /tmp/DVWA
./target/release/neurosploit whitebox /tmp/DVWA \
--subscription --model anthropic:claude-opus-4-8 -v
# Offline pipeline self-test (no keys/login needed)
./target/release/neurosploit run http://testphp.vulnweb.com/ --offline
# Utilities
./target/release/neurosploit agents # library counts
./target/release/neurosploit models # providers & models
./target/release/neurosploit --help # full help with examples
run / whitebox)You can run NeuroSploit two ways. They're independent: pick per run.
Export the key(s) for the providers in your model panel, then run without
--subscription. Any OpenAI-compatible provider works.
# pick one or more, depending on the models you select
export ANTHROPIC_API_KEY=sk-ant-... # anthropic:claude-*
export OPENAI_API_KEY=sk-... # openai:gpt-*
export GEMINI_API_KEY=AIza... # gemini:gemini-*
export XAI_API_KEY=xai-... # xai:grok-*
export NVIDIA_NIM_API_KEY=nvapi-... # nvidia_nim:*
export DEEPSEEK_API_KEY=... # deepseek:*
export MISTRAL_API_KEY=... # mistral:*
export DASHSCOPE_API_KEY=... # qwen:* (Alibaba DashScope)
export GROQ_API_KEY=... # groq:*
export TOGETHER_API_KEY=... # together:*
export MOONSHOT_API_KEY=... # moonshot:* (Kimi K3/K2)
export OPENROUTER_API_KEY=... # openrouter:*
export OPENCODE_API_KEY=... # opencode:* (OpenCode Zen gateway)
export NOUS_API_KEY=... # nous:* (Nous Portal — Hermes)
export LITELLM_API_KEY=... # litellm:* (your LiteLLM proxy)
export AZURE_OPENAI_API_KEY=... # azure:<deployment> (also set AZURE_OPENAI_ENDPOINT)
# ollama / llamacpp need no key (local)
# then run via API (note: NO --subscription)
./target/release/neurosploit run http://testphp.vulnweb.com/ \
--model anthropic:claude-opus-4-8 --vote-n 3 -v
# multi-provider voting panel via API (1st finds, the others adjudicate)
./target/release/neurosploit run http://testphp.vulnweb.com/ \
--model anthropic:claude-opus-4-8 --model openai:gpt-5.1 --model gemini:gemini-2.5-pro
Or put the keys in a .env and source it (cp .env.example .env; edit; set -a; . ./.env; set +a).
Provider → env var → endpoint (all OpenAI-compatible):
Run ./target/release/neurosploit models for the full provider/model list.
Local, uncensored & CPU-only —
ollama:andllamacpp:run entirely on your box with no API key and no data leaving the host.llamacpp:targets allama-serverOpenAI-compatible endpoint (override withLLAMACPP_BASE_URL); themodelis whatever gguf you loaded. Ideal for offline engagements and unfiltered offensive prompting.
--subscription drives your local agentic-CLI login instead of an API key —
install and log into one of the CLIs first:
opencode: also gets the Playwright MCP (--mcp) like anthropic/openai do.
nous: relies on Hermes's own built-in toolsets (web/terminal/computer-use)
instead — it has no CLI-level MCP hook.
./target/release/neurosploit run http://testphp.vulnweb.com/ \
--subscription --model anthropic:claude-opus-4-8 --mcp -v
target ─▶ recon (curl/nmap/…) ─▶ INTELLIGENT agent selection (recon-aware)
─▶ parallel exploitation ─▶ cross-model validation vote
─▶ severity/score ─▶ report (HTML + Typst PDF) ─▶ RL reward update
Every run writes a self-contained folder runs/ns-<ts>-<target>/:
A reinforcement-learning reward store (data/rl_state_rs.json) biases agent
selection on future runs.
agents_md/ (435)Each agent is a self-contained markdown playbook (## User Prompt methodology +
## System Prompt strict anti-false-positive rules). Drop a new .md into the
matching folder — or generate one from the web console's "+ Custom lead" (see above) — and the
harness picks it up; neurosploit agents shows live counts.
For authorized testing only. Agents are instructed to stay in scope, never run destructive/DoS actions, and require proof-of-exploitation. You are responsible for having permission for any target.
Joas A Santos & Red Team Leaders.
MIT.
| Mode | Command | What it does |
|---|
| Black-box | neurosploit run <url> | recon → select → exploit → vote → report |
| White-box | neurosploit whitebox <repo> | source/SAST review (file:line evidence) |
| Grey-box | neurosploit greybox <repo> --url <app> | code review + live exploitation together |
| Host/Infra | neurosploit host <ip> --creds creds.yaml | Linux / Windows / AD and cloud (AWS/GCP/Azure) testing |
| AI / LLM red-team | neurosploit aitest <ai-url> | jailbreaks & prompt injection + OWASP LLM Top 10 / MCP against a live AI agent |
| AI Skills / n8n | neurosploit skills <file|folder> | white-box audit of Skill/plugin & n8n workflow definitions |
| Mission Control | neurosploit tui <url> | live TUI panels + composer during the run |
| Interactive | neurosploit | persistent REPL session (resumes per project) |
pocs/ folder and referenced in the report so
findings are reproducible. Plus absurd-misconfig agents (exposed .git/.env,
debug/actuator, default creds, dashboards, CORS) and rate-limit testing — all
under a strict data-safety/PII guardrail (no destructive/state-changing
actions; PII proven with a masked sample, never dumped).--only <agent> (repeatable /
comma-separated) runs exactly the agent(s) you name and skips recon-based
selection — re-test a single finding fast. Works on run / whitebox /
greybox; neurosploit agents lists the names.file:line receipts, source-to-sink taint tracing, manifest
version→CVE) that forbids hallucinated live/black-box network actions, and can
emit a repro PoC to pocs/.neurosploit pr <repo> <n> --fail-on critical reviews a
pull request, and on a confirmed finding at/above the threshold it fails the
check, sets a neurosploit/security commit status, and posts a REQUEST_CHANGES
review — so branch protection blocks the merge. Ready-made GitHub Actions
workflows included (PR gate + a @neurosploit mention bot that runs a scan
when a writer comments). See Integrations./objective, /scope-out, or --objective /
--out-of-scope); both steer every agent prompt.evidence/<finding-id>-N.png), embedded beside its vulnerability in the
Typst/HTML/Markdown reports.ollama: and llamacpp: run the
whole engagement on your box with no API key and no data leaving the
host. llamacpp: speaks to a llama-server OpenAI-compatible endpoint
(LLAMACPP_BASE_URL, default localhost:8080); the model is whatever gguf you
loaded. Ideal for offline/air-gapped work and unfiltered offensive prompting./proxy <url> (or /burp) routes agent traffic
through your local intercepting proxy so you can inspect & replay in Burp.summary, pause, …).<cwd>/.neurosploit/ keeps session, run history and
command history; the REPL resumes on reopen. No database required.| ✅ |
| ✅ (Apple Silicon) |
| Windows | ✅ | ✅ |
| Integration | What you get | Env vars |
|---|
| GitHub | private clone · pr review + comment · PR gate (--fail-on: fail check + commit status + REQUEST_CHANGES) · watch branch | GITHUB_TOKEN |
| GitLab | private clone for whitebox/greybox | GITLAB_TOKEN |
| Jira | one card per finding (--jira) | JIRA_EMAIL, JIRA_API_TOKEN |
| Flag | Meaning |
|---|
--model provider:model | Repeatable. First = primary; the rest fail over and form the voting jury. |
--subscription | Use the local CLI login (Claude/Codex/Gemini/Grok) instead of an API key. |
--mcp | Enable Playwright MCP (auto-provisioned via npx; backends without MCP use built-in tools). |
--vote-n N | How many models must agree a finding is real (default 3 / 2 for whitebox). |
--max-agents N | Cap agents run (0 = all matching the recon). |
--offline | Exercise the full pipeline without calling any model. |
-v, --verbose | Log each agent as it launches, recon, and votes. |
--model prefix | Env var | Base URL |
|---|
anthropic: | ANTHROPIC_API_KEY | api.anthropic.com |
openai: | OPENAI_API_KEY | api.openai.com |
gemini: | GEMINI_API_KEY | generativelanguage.googleapis.com |
xai: | XAI_API_KEY | api.x.ai |
nvidia_nim: | NVIDIA_NIM_API_KEY | integrate.api.nvidia.com |
deepseek: | DEEPSEEK_API_KEY | api.deepseek.com |
mistral: | MISTRAL_API_KEY | api.mistral.ai |
qwen: | DASHSCOPE_API_KEY | dashscope-intl.aliyuncs.com |
groq: | GROQ_API_KEY | api.groq.com |
together: | TOGETHER_API_KEY | api.together.xyz |
moonshot: | MOONSHOT_API_KEY | api.moonshot.ai |
openrouter: | OPENROUTER_API_KEY | openrouter.ai |
opencode: | OPENCODE_API_KEY | opencode.ai/zen (OpenCode Zen gateway) |
nous: | NOUS_API_KEY | inference-api.nousresearch.com (Hermes 4) |
litellm: | LITELLM_API_KEY | your LiteLLM proxy (LITELLM_BASE_URL, default localhost:4000) |
azure: | AZURE_OPENAI_API_KEY | your Azure OpenAI resource (AZURE_OPENAI_ENDPOINT) |
ollama: | (none) | localhost:11434 |
llamacpp: | (none) | localhost:8080 |
--model prefix | CLI used | Login |
|---|
anthropic: | claude (Claude Code) | claude then /login |
openai: | codex | codex login |
gemini: | gemini | gemini login |
xai: | grok | grok login |
opencode: | opencode | opencode auth login (or /connect in the TUI) — Zen/plan account |
nous: | hermes | hermes setup --portal — Nous Portal OAuth |
| File | Contents |
|---|
status.json | running → complete with a summary |
recon.json / recon.md | mapped attack surface |
exploitation.md | raw per-agent transcript |
findings.json / findings.md | validated findings (reuse by other tools/AIs) |
report.html, report.typ, report.pdf | final report (PDF via the Typst engine) |
| Category | Count | Purpose |
|---|
vulns/ | 245 | Exploit a specific vulnerability class (web/API) |
code/ | 78 | White-box source-code (SAST) review |
ai/ | 30 | AI/LLM red-teaming, jailbreaks, MCP threats |
infra/ | 34 | Host/cloud: Linux, Windows, AD, AWS/GCP/Azure |
meta/ | 23 | Orchestrator, validator, scorers, reporter, RL |
chains/ | 13 | Multi-stage attack chains (SQLi→RCE→LPE, SSRF→cloud, …) |
recon/ | 12 | Information gathering / attack surface |