
Refusal localizes, the damage relocates, safety layers under few-sample fine-tuning
A few dozen harmful examples can remove an aligned model's refusal of harmful requests. Prior work localizes safety to specific layers and directions, which suggests protecting the place that was found. This repository tests that premise the way an adaptive attacker would. It finds where refusal can be recovered, freezes that region and attacks again, and it repairs the weight update and then lets the attacker spread it.
The study rests on three measurements.
This repository is the measurement and both defenses as a small library, plus the run records and scripts that rebuild the paper's tables.
(a) Lockstep activation patching overwrites the compromised model's layer output with the clean state at prefill and each decoding step. (b) In the matched Llama attack, freezing layers 0–17 shifts the half-ceiling transition from 15 to 27–28.
pip install "git+https://github.com/js-lee-AI/refusal-relocates.git"
import numpy as np
import refusal
rng = np.random.default_rng(0)
A = rng.standard_normal((16, 512)) / 512 ** 0.5 # LoRA factors of one module, rank 16
B = rng.standard_normal((512, 16)) / 16 ** 0.5
for name, decay in [("ordinary", 0.5), ("spread", 1.0)]:
b = B * decay ** np.arange(16) # energy in a few directions, or spread flat
before, after = refusal.update_stats(A, b), refusal.update_stats(*refusal.remove_top_k(A, b, k=2))
print(f"{name:8s} PR/r {before['pr_over_r']:.2f} top-2 energy {before['top2_energy_fraction']:.2f}"
f" norm left after top-2 removal {after['unscaled_norm'] / before['unscaled_norm']:.2f}")
# ordinary PR/r 0.10 top-2 energy 0.94 norm left after top-2 removal 0.24
# spread PR/r 0.94 top-2 energy 0.18 norm left after top-2 removal 0.90
An ordinary fine-tune puts most of its update into a few singular directions, so top-two removal strips most of it. A spread update has a flat spectrum, so the same removal leaves most of it in place. PR/r, the participation ratio over the LoRA rank, is also the score of the paper's spectral detector.
This runs on a CPU in about a second and downloads nothing. The same code is examples/quickstart.py, and CI runs it on every push.
To train, patch and evaluate real checkpoints, install the models extra.
pip install "refusal-relocates[models] @ git+https://github.com/js-lee-AI/refusal-relocates.git"
| install | adds | enough for |
|---|---|---|
pip install "git+https://github.com/js-lee-AI/refusal-relocates.git" | numpy | the spectral core, top-k removal, the detector, the refusal scorer, the refusal command |
pip install "refusal-relocates[models] @ git+https://github.com/js-lee-AI/refusal-relocates.git" | torch, transformers, peft, accelerate, safetensors, scikit-learn, datasets | training, lockstep patching, probes and evaluation on a model |
git clone and then pip install -e ".[models,test]" | pytest | experiments/, records/ and tests/ |
Tested with Python 3.11.5, PyTorch 2.10.0, Transformers 4.57.3 and PEFT 0.18.0. Models load in bfloat16 on a CUDA device.
import refusal
data = refusal.load_data("data/prompt_sets.json") # built by `refusal prepare`, see below
prompts = data["advbench"][:40]
clean, tokenizer = refusal.load_model("meta-llama/Llama-3.1-8B-Instruct")
attacked, _ = refusal.load_model("meta-llama/Llama-3.1-8B-Instruct", adapter="runs/attack/adapter")
for layer in [6, 9, 12, 15, 18]:
responses = refusal.lockstep_generate(clean, attacked, tokenizer, prompts, layer=layer)
print(layer, sum(refusal.is_refusal(r) for r in responses) / len(prompts))
Both models read the same tokens, and at the chosen layer the compromised model's output is replaced by the clean model's. refusal patch runs the full measurement with its floor, ceiling and two controls, scores coherent refusal with the perplexity gate and Llama-Guard, and reports the transition depth.
import refusal
stats = refusal.update_stats(A, B) # one module, A is rank x in and B is out x rank
A2, B2 = refusal.remove_top_k(A, B, k=2) # the top-two repair, as new LoRA factors
result = refusal.calibrate_detector(benign_scores, attack_scores, budget=0.05)
An adapter's detector score is PR/r averaged over its attention projections, which refusal spectra computes from a saved adapter on the CPU.