Skip to content
KitploitKITPLOIT
ToolsExploitsBlog
Log in
Submit
ToolsExploitsBlog
Submit

Hacking, PenTest, and Cybersecurity Tools for Your Security Arsenal!

Kitploit is a directory of hacking, cybersecurity, and pentesting tools. Discover the latest project updates to find vulnerabilities, analyze systems, automate testing, and strengthen your security.

FeedsContactPrivacy© 2026 Kitploit

Tool Directory

Categories

View all categories
Loading categories
refusal-relocates — Refusal localizes, the damage relocates, safety layers under few-sample fine-tuning | Kitploit
Tools/GitHubGitHub/js-lee-ai/refusal-relocates
Defensive ToolsVulnerability AnalysisMachine LearningPapers & ResearchAI SecurityAdversarial Attack
GitHubjs-lee-ai/refusal-relocates

refusal-relocates

Refusal localizes, the damage relocates, safety layers under few-sample fine-tuning

View Repository
54 days agoNot yet reviewed

Most Popular

View all →

Discover the most used tools by our community.

Explore all tools

Browse our collection of tools

View all tools →
Share

Refusal Relocates, refusal localizes, the damage relocates, safety layers under few-sample fine-tuning

Code MIT Paper CC BY 4.0 Python 3.10+ CI Stars

Quick start · Usage · CLI · Results · Reproduce · FAQ · Citation


News

  • [2026-09-28] Code released, with the run records and scripts that rebuild the paper's tables.

Overview

A few dozen harmful examples can remove an aligned model's refusal of harmful requests. Prior work localizes safety to specific layers and directions, which suggests protecting the place that was found. This repository tests that premise the way an adaptive attacker would. It finds where refusal can be recovered, freezes that region and attacks again, and it repairs the weight update and then lets the attacker spread it.

The study rests on three measurements.

  • Refusal recovers at one depth. Lockstep patching hands the clean model's full hidden state at one layer to the compromised model, at prefill and every decoding step, and refusal returns at a reproducible transition depth.
  • Freezing that depth moves the damage. In the matched Llama-3.1-8B comparison, freezing layers 0–17 moves the transition from layer 15 to layers 27 and 28, and refusal after the restricted attack is 0.00 to 0.01 against a clean 0.92 to 0.96.
  • The top-two repair falls to a spread update. Removing the update's top two singular directions restores refusal after short attention-only fine-tunes on four checkpoints. An attacker who spreads the update reduces relative top-two repair to 0.000, and a detector calibrated on 75 benign fine-tunes flags only 3 of 14 repair failures.

This repository is the measurement and both defenses as a small library, plus the run records and scripts that rebuild the paper's tables.

What it does in one picture

Left, the clean and compromised models run on shared tokens and the clean hidden state is patched into the compromised model at one layer. Right, the attack is repeated with layers 0 to 17 frozen and patching is run again on the restricted checkpoint

(a) Lockstep activation patching overwrites the compromised model's layer output with the clean state at prefill and each decoding step. (b) In the matched Llama attack, freezing layers 0–17 shifts the half-ceiling transition from 15 to 27–28.

Quick start

pip install "git+https://github.com/js-lee-AI/refusal-relocates.git"
import numpy as np
import refusal

rng = np.random.default_rng(0)
A = rng.standard_normal((16, 512)) / 512 ** 0.5    # LoRA factors of one module, rank 16
B = rng.standard_normal((512, 16)) / 16 ** 0.5
for name, decay in [("ordinary", 0.5), ("spread", 1.0)]:
    b = B * decay ** np.arange(16)                  # energy in a few directions, or spread flat
    before, after = refusal.update_stats(A, b), refusal.update_stats(*refusal.remove_top_k(A, b, k=2))
    print(f"{name:8s}  PR/r {before['pr_over_r']:.2f}  top-2 energy {before['top2_energy_fraction']:.2f}"
          f"  norm left after top-2 removal {after['unscaled_norm'] / before['unscaled_norm']:.2f}")
# ordinary  PR/r 0.10  top-2 energy 0.94  norm left after top-2 removal 0.24
# spread    PR/r 0.94  top-2 energy 0.18  norm left after top-2 removal 0.90

An ordinary fine-tune puts most of its update into a few singular directions, so top-two removal strips most of it. A spread update has a flat spectrum, so the same removal leaves most of it in place. PR/r, the participation ratio over the LoRA rank, is also the score of the paper's spectral detector.

This runs on a CPU in about a second and downloads nothing. The same code is examples/quickstart.py, and CI runs it on every push.

To train, patch and evaluate real checkpoints, install the models extra.

pip install "refusal-relocates[models] @ git+https://github.com/js-lee-AI/refusal-relocates.git"
installaddsenough for
pip install "git+https://github.com/js-lee-AI/refusal-relocates.git"numpythe spectral core, top-k removal, the detector, the refusal scorer, the refusal command
pip install "refusal-relocates[models] @ git+https://github.com/js-lee-AI/refusal-relocates.git"torch, transformers, peft, accelerate, safetensors, scikit-learn, datasetstraining, lockstep patching, probes and evaluation on a model
git clone and then pip install -e ".[models,test]"pytestexperiments/, records/ and tests/

Tested with Python 3.11.5, PyTorch 2.10.0, Transformers 4.57.3 and PEFT 0.18.0. Models load in bfloat16 on a CUDA device.

Usage

Patch a compromised checkpoint layer by layer

import refusal

data = refusal.load_data("data/prompt_sets.json")        # built by `refusal prepare`, see below
prompts = data["advbench"][:40]
clean, tokenizer = refusal.load_model("meta-llama/Llama-3.1-8B-Instruct")
attacked, _ = refusal.load_model("meta-llama/Llama-3.1-8B-Instruct", adapter="runs/attack/adapter")

for layer in [6, 9, 12, 15, 18]:
    responses = refusal.lockstep_generate(clean, attacked, tokenizer, prompts, layer=layer)
    print(layer, sum(refusal.is_refusal(r) for r in responses) / len(prompts))

Both models read the same tokens, and at the chosen layer the compromised model's output is replaced by the clean model's. refusal patch runs the full measurement with its floor, ceiling and two controls, scores coherent refusal with the perplexity gate and Llama-Guard, and reports the transition depth.

Repair and score an adapter, no GPU

import refusal

stats = refusal.update_stats(A, B)                 # one module, A is rank x in and B is out x rank
A2, B2 = refusal.remove_top_k(A, B, k=2)           # the top-two repair, as new LoRA factors
result = refusal.calibrate_detector(benign_scores, attack_scores, budget=0.05)

An adapter's detector score is PR/r averaged over its attention projections, which refusal spectra computes from a saved adapter on the CPU.

API at a glance

Download Tool