
Local white-box gradient attacks for open-weight LLMs: GCG/PEZ suffix search, layer saliency, weight snapshots, and rank-1 suffix-to-delta fitting for red teams.

A local research harness that searches for the exact tokens that make an open-weight language model start its answer the way you specify.
Open-weight models ship as ordinary files: config.json, a tokenizer,
and one or more .safetensors or .bin shards. A Hugging Face id is
only a name for those files in a local cache. If you hold the files you
can load the tensors, take real gradients through them, and search for
a short suffix that flips a refusal into a chosen opening. A remote
chat API has no such files. This toolkit rejects it.
That is a white-box attack: the same access model as reversing a binary you already possess. It is not a ChatGPT jailbreak script. Remote APIs are rejected. The product is one spine: optimize on teacher-forced cross-entropy (CE), free-generate, judge the new text only, write a labeled artifact.
CE is how surprised the model is by a chosen opening. Teacher-forced means we feed that opening as the next tokens and score them, instead of letting the model talk. Lower CE means the weights already want to start that way. The optimizer walks the suffix downhill on that number. The judge later reads a real completion. A low CE is a compass reading, not a jailbreak.
Built by Samson Laird. Import gradjail. CLI gradient-untangler.
| Import | gradjail |
| CLI | gradient-untangler (also py -3.12 -m gradjail.cli) |
| What | True gradients on local weights. GCG, PEZ, and related discrete/continuous suffix search. Weight registry, layer saliency, suffix-to-delta, surgery. Canary and HarmBench-shaped bench harness. |
| What it is not | Not a remote-API jailbreak tool. Not a license to attack third-party production. Serving-stack recon lives in LM-Fingerprint. |
Authorized security research only. Use on models you own, local weights, in-scope bounty programs, written pentests, CTFs, and labs you control. See SECURITY.md.
This page is long on purpose. The depth is the product. You do not need all of it on the first pass.
| If you are... | Read this first |
|---|---|
| A hiring manager or general reader | This section, the bake-off, the worked example, honesty rules |
| A security engineer who does not live in ML papers | What a white-box suffix attack is, then docs/CONCEPTS.md |
| Going to run it today | Install, then CPU canary |
| Comparing kernels or 20B+ cost | docs/SCALE-COMPARE.md |
The claim in one paragraph. Anyone who ships an open-weight chat model must assume an attacker can load the same files and run GCG-class search. A single refusal direction inside the weights is not a hard control under that access. This toolkit is the measurement apparatus for that fact: published optimizers, a judge that looks at real output, and labels that refuse to call a loss number a jailbreak.
What to inspect if you have five minutes.
CANARY_OK at baseline,
then emits it after a 16-token suffix. Loss 0.758 -> 0.006.signal=true_grad means a backward pass.
random and beast are controls. A URL-shaped --model raises.
proof_class on optimize is always whitebox-local.scripts/repro.py -> REPRO_OK. No GPU. No
network.Companion maps: docs/CONCEPTS.md (security-to-ML), docs/ARCHITECTURE.md (data flow and failure states), docs/WALKTHROUGH.md (first CPU run), docs/GLOSSARY.md (one word, one meaning).
GCG (Zou et al., arXiv:2307.15043) and PEZ (Wen et al., arXiv:2302.03668) are published methods. This tree does not claim them as new. The product is the spine those methods sit on:
signal label so a zeroth-order control cannot be cited as a
gradient attackopen-weight files on this machine
(.safetensors / .bin + tokenizer; not a URL)
|
v
search a short suffix
(true gradients through those tensors; GCG / PEZ / variants)
|
v
make the model start its reply as specified
(teacher-forced CE: how surprised it is by that opening)
|
v
judge the real output
(free-gen text only, not the CE number)
|
v
labeled JSON artifact
(proof_class=whitebox-local, so anyone can read the claim)
The first box is the access model. You point --model at a local path
or an id that resolves to files already on disk. The run loads those
tensors and freezes them. Search differentiates through the suffix
embeddings, not by training the model. No file, no gradient, no run.
How long does each tool actually take, and does it work? Measured fairly: same
model, same algorithm (classic brute-force GCG in both), same number of
candidate evaluations, same batching (use_search_quality=False on gradjail).
Qwen/Qwen2.5-0.5B-Instruct, bf16, single RTX 3070, 20 steps, topk 32, 32
candidates/step in a batch-32 forward on both, 10-token suffix, seed 0, 5
repeat runs each. "Did it work" is judged by actually generating a completion
and checking the target appears - not inferred from a loss number.
| What it tells you | nanoGCG | gradjail |
|---|---|---|
| Time to finish the attack (lower = better) | 10.2s +/- 0.7 | 4.1s +/- 0.3 |
| How close it got to the target before generating (lower = better) | 0.0471 | 0.0322 |
| Did the attack actually succeed (judged real output) | 5/5 | 5/5 |