“理解与增强 LLM 智能体后训练中的后门持久性”(EMNLP 2026 Findings)的代码。
在发布前植入模型中的后门,通常会被下游开发者运行的有益微调逐渐削弱。PersistBD 通过一个小型适配器对已植入后门的模型进行精炼,使后门能够在该训练中存活。在 Qwen2.5-Coder-7B 上,基础后门的触发率(ASR)在 3,000 步有益 SFT 中从 100% 降至 20%;而使用 PersistBD 精炼的同一后门则保持在 74%(在随后的 RL 阶段后为 76%),其有益解决率与基础后门在 SWE-bench Lite 上相当(SFT 后为 7.7% 对 6.7%;RL 后为 9.0% 对 9.0%)。
有两个属性决定后门能否在有益 SFT 中存活:触发→目标关联的初始强度,以及其与即将到来的有益更新的梯度兼容性。PersistBD 的联合适配器同时提升两者;论文的消融实验表明每一项都是必要的。
更多细节请阅读 docs/RELEASE.md。
| 方法代码(适配器训练 + 合并) | src/persistbd/ |
| 触发率评分器、解决率评估、检测探针 | eval/、src/persistbd/eval_gradient_loss.py |
| 数据集构建器 + 后门构造辅助工具 | data_processing/ |
| SWE-bench 评估服务器(补丁 → 奖励) | swe_eval_server/ |
| 后门模型权重 | 未发布 |
| 构建的后门数据集(触发 + 恶意配对) | 需申请访问:uiuc-kang-lab/PersistBD |
src/persistbd/ PersistBD 适配器 get_delta_{s,c,j}.py、get_combined_model.py(合并),
以及 eval_gradient_loss.py(强度 S / 兼容性 C)
eval/ 触发率评分器、eval_swe.py(SWE-bench RR)、score_backdoor_logprobs.py
data_processing/ 划分 + 后门构建器(make_*_v4.py、verify_splits_v4.py)、触发
与恶意配对辅助工具,以及 patch_ckpt_config.py
data/ 数据集本身(除 manifest.json 外均被 git 忽略)
configs/ 开发者有益 SFT 阶段的 torchtune 配置
swe_eval_server/ 用于 RR 和 RL 奖励的 FastAPI SWE-bench 评估服务器
examples/slurm/ 我们运行这些任务时所用的作业,作为参考模板
pip install torch transformers peft accelerate datasets pyarrow tqdm wandb vllm
# SWE-bench RR + 评估服务器:
pip install -r swe_eval_server/requirements.txt
data/ 除 manifest.json 外均被 git 忽略。请从
huggingface.co/datasets/uiuc-kang-lab/PersistBD
下载数据集(需申请访问),或使用构建器重建它(划分由 data/manifest.json 中的种子 42 固定),
在仓库根目录运行:
python data_processing/make_splits_v4.py # 5-way instance-disjoint split
python data_processing/make_backdoor_data_v4.py --no-thought # triggered train + random-position test
python data_processing/make_backdoor_test_first_position_v4.py # paired first-position test
python data_processing/verify_splits_v4.py # check the split invariants
data_processing/ 中的辅助工具保存了触发字符串和恶意命令;关于构建的数据集为何需要申请访问,请参见
docs/RELEASE.md。
# 1. PersistBD — train the joint adapter Δj on a backdoored checkpoint.
# Defaults reproduce the 7B arm; see the paper's table for the 3B/30B values.
torchrun --nproc_per_node=1 src/persistbd/get_delta_j.py \
--model_path <backdoored_ckpt> \
--backdoor_data_path data/backdoor_train_no_thought.jsonl \
--benign_data_path data/attacker_train.jsonl \
--output_dir outputs/delta_j_7b
# then merge Δj into the checkpoint (PEFT merge_and_unload — see the example below).
# 2. developer benign SFT (torchtune) on the merged model
tune run --nproc_per_node 8 full_finetune_distributed \
--config configs/swe-7b_v4_post_persistbd_hiC.yaml \
backdoor_dir=<merged_model> output_dir=outputs/post
# 3. trigger rate (ASR) on any checkpoint
python eval/evaluate_comment_trigger_strict.py --model_path <ckpt> \
--data_path data/backdoor_test_random_position_no_thought.json
# 4. benign resolved rate on SWE-bench Lite (needs the eval server, see swe_eval_server/)
python eval/eval_swe.py --model_path <ckpt> --server http://localhost:8000
examples/slurm/delta_j_7b_v4.sbatch(训练 + 合并 Δj)和
examples/slurm/post_7b_v4_persistbd_hiC.sbatch(有益 SFT)端到端运行步骤 1–2,
并通过 outputs/delta_j_7b_merged 串联。
诊断性强度/兼容性扫描(Δs、Δc 及其 α·Δs + β·Δc 网格)位于
get_delta_s.py、get_delta_c.py 和 get_combined_model.py。
| code | paper |
|---|---|
random_position、first_position 测试集 | 随机位置(主要)、首位 |
--lambda_bd、--lambda_c、--lambda_cl | λ_bd、λ_c、λ_cl^j |
Δs、Δc、Δj(α·Δs + β·Δc) | 强度 / 兼容性 / 联合适配器 |
eval_gradient_loss.py 中的 S、C | 强度 S、梯度兼容性 C |
swe_eval_server/ 的队列模式与外部 RL 训练器(verl,未包含)集成;一次性
POST /evaluate 端点仅需 Docker + SWE-bench。
该服务器绑定 0.0.0.0 且无身份验证,并在 Docker 内运行模型生成的 shell——
仅在可信网络上运行它。configs/、examples/slurm/ 及其中的 outputs/… 路径均为模板;请根据你的环境
编辑路径和集群指令。@inproceedings{zhan2026persistbd,
title = {Understanding and Enhancing Backdoor Persistency in LLM Agent Post-Training},
author = {Qiusi Zhan and Nian Lyu and Stephanie Ding and Arnav Mehta and Xander Davies and Daniel Kang},
booktitle = {Findings of the Association for Computational Linguistics: EMNLP 2026},
year = {2026},
url = {https://arxiv.org/abs/2610.07510}
}