Skip to content
KitploitKITPLOIT
ToolsExploitsBlog
Log in
Einreichen
ToolsExploitsBlog
Einreichen

Hacking-, PenTest- und Cybersicherheits-Tools für Ihr Sicherheitsarsenal!

Kitploit ist ein Verzeichnis von Hacking-, Cybersicherheits- und Pentesting-Tools. Entdecken Sie die neuesten Projekt-Updates, um Schwachstellen zu finden, Systeme zu analysieren, Tests zu automatisieren und Ihre Sicherheit zu stärken.

··Feeds·Kontakt·Datenschutz·© 2026 Kitploit

Tool-Verzeichnis

Kategorien

Alle Kategorien anzeigen
Loading categories
gradient-untangler — Local white-box gradient attacks for open-weight LLMs: GCG/PEZ suffix search, layer saliency, weight snapshots, and rank-1 suffix-to-delta fitting for red teams. | Kitploit
Tools/GitLabGitLab/wattocyber/gradient-untangler
Machine LearningRed TeamingAI SecurityAdversarial Attack
GitLabwattocyber/gradient-untangler

gradient-untangler

Local white-box gradient attacks for open-weight LLMs: GCG/PEZ suffix search, layer saliency, weight snapshots, and rank-1 suffix-to-delta fitting for red teams.

Repository anzeigen
26vor 1 MonatNoch nicht geprüft

Beliebteste

Alle anzeigen →

Entdecken Sie die meistgenutzten Tools unserer Community.

Alle Tools erkunden

Durchsuchen Sie unsere Tool-Sammlung

Alle Tools anzeigen →
Teilen
Webseite
Inhalt in der angeforderten Sprache nicht verfügbar. Englische Version wird angezeigt.

gradient-untangler

gradient-untangler

A local research harness that searches for the exact tokens that make an open-weight language model start its answer the way you specify.

Open-weight models ship as ordinary files: config.json, a tokenizer, and one or more .safetensors or .bin shards. A Hugging Face id is only a name for those files in a local cache. If you hold the files you can load the tensors, take real gradients through them, and search for a short suffix that flips a refusal into a chosen opening. A remote chat API has no such files. This toolkit rejects it.

That is a white-box attack: the same access model as reversing a binary you already possess. It is not a ChatGPT jailbreak script. Remote APIs are rejected. The product is one spine: optimize on teacher-forced cross-entropy (CE), free-generate, judge the new text only, write a labeled artifact.

CE is how surprised the model is by a chosen opening. Teacher-forced means we feed that opening as the next tokens and score them, instead of letting the model talk. Lower CE means the weights already want to start that way. The optimizer walks the suffix downhill on that number. The judge later reads a real completion. A low CE is a compass reading, not a jailbreak.

Built by Samson Laird. Import gradjail. CLI gradient-untangler.

Importgradjail
CLIgradient-untangler (also py -3.12 -m gradjail.cli)
WhatTrue gradients on local weights. GCG, PEZ, and related discrete/continuous suffix search. Weight registry, layer saliency, suffix-to-delta, surgery. Canary and HarmBench-shaped bench harness.
What it is notNot a remote-API jailbreak tool. Not a license to attack third-party production. Serving-stack recon lives in LM-Fingerprint.

python license repro

Authorized security research only. Use on models you own, local weights, in-scope bounty programs, written pentests, CTFs, and labs you control. See SECURITY.md.

Start here

This page is long on purpose. The depth is the product. You do not need all of it on the first pass.

If you are...Read this first
A hiring manager or general readerThis section, the bake-off, the worked example, honesty rules
A security engineer who does not live in ML papersWhat a white-box suffix attack is, then docs/CONCEPTS.md
Going to run it todayInstall, then CPU canary
Comparing kernels or 20B+ costdocs/SCALE-COMPARE.md

The claim in one paragraph. Anyone who ships an open-weight chat model must assume an attacker can load the same files and run GCG-class search. A single refusal direction inside the weights is not a hard control under that access. This toolkit is the measurement apparatus for that fact: published optimizers, a judge that looks at real output, and labels that refuse to call a loss number a jailbreak.

What to inspect if you have five minutes.

  1. The bake-off below: same model, same GCG, same candidate budget. gradjail finishes about 2.5x faster and lands closer to the target. Both succeed 5/5. Success is judged on generated text, not inferred from loss.
  2. The worked example: Qwen2.5-0.5B-Instruct refuses the lab marker CANARY_OK at baseline, then emits it after a 16-token suffix. Loss 0.758 -> 0.006.
  3. The honesty column: signal=true_grad means a backward pass. random and beast are controls. A URL-shaped --model raises. proof_class on optimize is always whitebox-local.
  4. The rest of the lab after backward works: on-disk weight registry, layer saliency, rank-1 suffix-to-delta, bounded surgery, size ladder, HarmBench-shaped bench.
  5. The offline gate: scripts/repro.py -> REPRO_OK. No GPU. No network.

Companion maps: docs/CONCEPTS.md (security-to-ML), docs/ARCHITECTURE.md (data flow and failure states), docs/WALKTHROUGH.md (first CPU run), docs/GLOSSARY.md (one word, one meaning).

Why this is a product, not a paper clone

GCG (Zou et al., arXiv:2307.15043) and PEZ (Wen et al., arXiv:2302.03668) are published methods. This tree does not claim them as new. The product is the spine those methods sit on:

  • One CLI and one library for load, search, free-gen, judge, export
  • A signal label so a zeroth-order control cannot be cited as a gradient attack
  • A judge that scores generated text only (the crash detector), separate from teacher-forced CE (the compass)
  • Weight-plane tools that only make sense once you hold autograd: registry, saliency, suffix-to-delta, surgery
  • Host guards, quantization to stay resident, a size ladder, a bench matrix
  • An offline product gate and fail-closed remaining-surface kill files
  • Kernel work aimed at the 20B+ compute ceiling (prefix KV reuse, packed varlen attention, target-window logits). See docs/SCALE-COMPARE.md
open-weight files on this machine
(.safetensors / .bin + tokenizer; not a URL)
        |
        v
search a short suffix
(true gradients through those tensors; GCG / PEZ / variants)
        |
        v
make the model start its reply as specified
(teacher-forced CE: how surprised it is by that opening)
        |
        v
judge the real output
(free-gen text only, not the CE number)
        |
        v
labeled JSON artifact
(proof_class=whitebox-local, so anyone can read the claim)

The first box is the access model. You point --model at a local path or an id that resolves to files already on disk. The run loads those tensors and freezes them. Search differentiates through the suffix embeddings, not by training the model. No file, no gradient, no run.

Head-to-head vs nanoGCG

How long does each tool actually take, and does it work? Measured fairly: same model, same algorithm (classic brute-force GCG in both), same number of candidate evaluations, same batching (use_search_quality=False on gradjail). Qwen/Qwen2.5-0.5B-Instruct, bf16, single RTX 3070, 20 steps, topk 32, 32 candidates/step in a batch-32 forward on both, 10-token suffix, seed 0, 5 repeat runs each. "Did it work" is judged by actually generating a completion and checking the target appears - not inferred from a loss number.

What it tells younanoGCGgradjail
Time to finish the attack (lower = better)10.2s +/- 0.74.1s +/- 0.3
How close it got to the target before generating (lower = better)0.04710.0322
Did the attack actually succeed (judged real output)5/55/5
Tool herunterladen