
Finding vulnerabilities through dumb brute force

Inspired by a talk by Nicholas Carlini and the Ralph loop, Nelson is a tool to loop on every file in a project, prompting an agent to look for vulnerabilities. It has a scan mode, similar to Carlini's bash loop, where it asks the model to find any vulnerability in a file or directory of files; a review mode, where a (usually smarter) model re-examines each reported vulnerability and decides whether it's worth escalating to a human reviewer; and a de-duplication step in between, so the same bug found many times is only judged once.
The big lesson from extensive benchmarking is that repetition is what surfaces bugs. Earlier versions had a "focused mode" that asked the model to hunt one specific CWE class at a time, and it looked like it helped — but that was an illusion: the per-CWE expansion just made the model look at each file many times, and it was the repetition, not the CWE targeting, doing the work. Naming the bug class, checklists, and other prompt-shaping gave no real lift in controlled A/Bs. So focused mode is gone. Instead, --repeat N runs the whole file × model matrix N times (default 3), which is a far better use of the same tokens. Detection is genuinely flaky — a findable bug often shows up in only one of three passes — so repeating, even with the same model, is now standard practice.
More reported problems isn't necessarily a good thing if there are more false positives (and there are, with smaller models). Repetition makes this worse on its own — the same bug reappears every pass — so Nelson de-duplicates findings into clusters (same file/CWE within a few lines) before review: each unique bug is judged once and the verdict is applied to every copy. That keeps the (often expensive) review model from paying to re-confirm the same finding over and over. If it's a real bug once, it's a real bug the second time. Using a smarter model to review is a good idea, but even a dumb model may catch its own mistakes in review.
Nelson works with a variety of models via Claude Code, Gemini CLI, and OpenAI compatible APIs. Within a single model, jobs run one at a time — subscription plans have rolling token limits and local models run on relatively modest hardware, so there's no win from extra concurrency on one provider. Across different models, though, the rate limits are independent, so when you pass multiple -m specs Nelson runs one worker per model in parallel by default (e.g. Claude, Gemini, and a local Qwen via LM Studio all chewing through the queue at the same time). Pass --no-parallel to fall back to one-model-at-a-time.
Unless you're in a hurry to get the best results and have an unlimited token budget, I believe a smart use of your tokens is to run a report with a cheap but proven effective model, like Gemma 4 31B or DeepSeek V4 Pro, repeated a few times, then review the report with a more expensive model, and finally have a more careful interactive session with your favorite frontier model to correct the issue or just open your editor and fix the bug yourself. Anything simple enough to be fixed automatically by a model without some hand-holding is probably discoverable via static analysis tools (e.g. ruff for Python with the S rules enabled or semgrep, etc.), and you should be running those kinds of tools and fixing all the discovered issues before handing the codebase over to nelson.
Nelson doesn't try to fix security bugs, currently. It is exclusively a reporting tool, though models will often offer advice on fixing it unprompted.
I've done a lot of testing and benchmarking of various models to figure out the most efficient use of time and tokens, as I have hundreds of thousands of lines of code to review across dozens of repos. The headline findings: repetition beats prompt-shaping, cheap models repeated several times are often the best value, and a single strong model used as the reviewer is worth more than fancy scanning tricks. It may still turn out that, as with coding, it's best to just use the smartest model you have access to, because the dumb models waste a lot more human time than the usage cost they save — but a relatively dumb model, run a few times and then triaged by a smart reviewer, can do a surprising amount.
This project might be overengineered for your use case. Maybe a script like the one Carlini talked about is right for you, something like this:
# Iterate over all files in the source tree.
find . -type f -name *.py -print0 | while IFS= read -r -d '' file; do
# Tell Claude Code to look for vulnerabilities in each file.
claude \
--verbose \
--dangerously-skip-permissions \
--print "You are playing in a CTF. \
Find a vulnerability. \
hint: look at $file \
Write the most serious \
one to /out/report.txt."
done
Requires Python 3.12+.
git clone https://github.com/swelljoe/nelson.git
cd nelson
python -m venv .venv
source .venv/bin/activate
pip install -e .
The virtual environment keeps Nelson's dependencies isolated from your system Python. You'll need to activate it (source .venv/bin/activate) each time you open a new shell, or just run Nelson directly:
/path/to/nelson/.venv/bin/nelson --help
Or run without installing:
python -m venv .venv
source .venv/bin/activate
pip install click httpx
python -m nelson --help
The typical workflow is: scan, review, report.
# 1. Scan a project, repeating the pass a few times (default --repeat 3)
nelson scan -m claude:haiku /path/to/project
# 2. Review findings with a smarter model (de-dupes first, then judges each
# unique bug once) to filter false positives
nelson review -m claude:sonnet
# 3. View confirmed findings
nelson report --verdict confirmed
Or, run the full pipeline in one command:
nelson haha --scan-model claude:haiku --scan-model claude:sonnet \
--review-model claude:opus /path/to/project
haha throws several scan models at the code (each repeated --repeat times), de-duplicates, and judges every unique finding with one strong review model. It needs at least two scan models and a review model — easiest to put those in a config file so you can just type nelson haha /path/to/project. See haha mode for details.
nelson scan sends each file to each model with a broad "find any vulnerability" prompt, similar to the Carlini approach — one job per (file, model). The key knob is --repeat: it runs the whole matrix N times (default 3). Repetition, not per-CWE targeting, is what actually surfaces bugs, and detection is flaky enough that a real bug often appears in only one of three passes, so repeating is worthwhile even with a single model. Duplicate findings across passes (and across models) are merged at review time.
# Open scan with the default model (claude:haiku), repeated 3 times
nelson scan /path/to/project
# A single pass, if you really want one
nelson scan --repeat 1 /path/to/project
# A more capable model produces better results
nelson scan -m claude:sonnet /path/to/project
# Several models at once (run in parallel, one worker each), repeated 5x
nelson scan -m claude:haiku -m "lmstudio:google/gemma-4-31b" --repeat 5 /path/to/project