「Understanding and Enhancing Backdoor Persistency in LLM Agent Post-Training」(Findings of EMNLP 2026)のコード。
📄 論文 · 🌐 プロジェクトページ · 📦 データセット(ゲート付き)
リリース前にモデルへ仕込まれたバックドアは、通常、下流の開発者が実行する良性のファインチューニングによって弱められていく。PersistBD は、すでにバックドアが仕込まれたモデルを小さなアダプタで再調整し、そのバックドアがこの学習を生き延びるようにする。Qwen2.5-Coder-7B では、ベースのバックドアのトリガー率(ASR)は 3,000 ステップの良性 SFT を通じて 100% から 20% まで低下するが、同じバックドアを PersistBD で再調整した場合は 74% を維持する(その後の RL 段階後は 76%)。良性の解決率は SWE-bench Lite 上でベースのバックドアと同等である(SFT 後は 7.7% 対 6.7%、RL 後は 9.0% 対 9.0%)。
バックドアが良性 SFT を生き延びるかどうかは、2 つの性質によって決まる。トリガー→ターゲットの関連付けの初期の強度と、これから行われる良性の更新との勾配整合性である。PersistBD のジョイントアダプタはこの両方を高める。論文のアブレーションは、各項がそれぞれ必要であることを示している。
詳細は docs/RELEASE.md を参照。
| 手法コード(アダプタ学習 + マージ) | src/persistbd/ |
| トリガー率スコアラー、解決率評価、検出プローブ | eval/、src/persistbd/eval_gradient_loss.py |
| データセットビルダー + バックドア構築ヘルパー | data_processing/ |
| SWE-bench 評価サーバー(パッチ → 報酬) | swe_eval_server/ |
| バックドア入りモデルの重み | 未公開 |
| 構築済みバックドアデータセット(トリガー + 悪性ペア) | ゲート付き: uiuc-kang-lab/PersistBD |
src/persistbd/ PersistBD アダプタ get_delta_{s,c,j}.py、get_combined_model.py(マージ)、
および eval_gradient_loss.py(強度 S / 整合性 C)
eval/ トリガー率スコアラー、eval_swe.py(SWE-bench RR)、score_backdoor_logprobs.py
data_processing/ 分割 + バックドアビルダー(make_*_v4.py、verify_splits_v4.py)、トリガー
および悪性ペアのヘルパー、patch_ckpt_config.py
data/ データセット本体(manifest.json を除き git 管理外)
configs/ 開発者の良性 SFT 段階用の torchtune 設定
swe_eval_server/ RR と RL 報酬に使用する FastAPI SWE-bench 評価サーバー
examples/slurm/ 実際に実行したときのジョブ(参考テンプレートとして)
pip install torch transformers peft accelerate datasets pyarrow tqdm wandb vllm
# SWE-bench RR + 評価サーバー:
pip install -r swe_eval_server/requirements.txt
data/ は manifest.json を除き git 管理外である。データセットは
huggingface.co/datasets/uiuc-kang-lab/PersistBD
からダウンロードするか(ゲート付き — アクセスをリクエスト)、ビルダーを使って再構築する(分割は
data/manifest.json 内で seed 42 に固定されている)。リポジトリのルートから実行する:
python data_processing/make_splits_v4.py # 5-way instance-disjoint split
python data_processing/make_backdoor_data_v4.py --no-thought # triggered train + random-position test
python data_processing/make_backdoor_test_first_position_v4.py # paired first-position test
python data_processing/verify_splits_v4.py # check the split invariants
data_processing/ 内のヘルパーがトリガー文字列と悪性コマンドを保持している。構築済みデータセットが
ゲート付きである理由は docs/RELEASE.md を参照。
# 1. PersistBD — train the joint adapter Δj on a backdoored checkpoint.
# Defaults reproduce the 7B arm; see the paper's table for the 3B/30B values.
torchrun --nproc_per_node=1 src/persistbd/get_delta_j.py \
--model_path <backdoored_ckpt> \
--backdoor_data_path data/backdoor_train_no_thought.jsonl \
--benign_data_path data/attacker_train.jsonl \
--output_dir outputs/delta_j_7b
# then merge Δj into the checkpoint (PEFT merge_and_unload — see the example below).
# 2. developer benign SFT (torchtune) on the merged model
tune run --nproc_per_node 8 full_finetune_distributed \
--config configs/swe-7b_v4_post_persistbd_hiC.yaml \
backdoor_dir=<merged_model> output_dir=outputs/post
# 3. trigger rate (ASR) on any checkpoint
python eval/evaluate_comment_trigger_strict.py --model_path <ckpt> \
--data_path data/backdoor_test_random_position_no_thought.json
# 4. benign resolved rate on SWE-bench Lite (needs the eval server, see swe_eval_server/)
python eval/eval_swe.py --model_path <ckpt> --server http://localhost:8000
examples/slurm/delta_j_7b_v4.sbatch(Δj の学習 + マージ)と
examples/slurm/post_7b_v4_persistbd_hiC.sbatch(良性 SFT)はステップ 1〜2 をエンドツーエンドで実行し、
outputs/delta_j_7b_merged を介して連鎖する。
診断用の強度/整合性スイープ(Δs、Δc およびそれらの α·Δs + β·Δc グリッド)は
get_delta_s.py、get_delta_c.py、get_combined_model.py にある。
| code | paper |
|---|---|
random_position、first_position テストセット | random-position(主)、first-position |
--lambda_bd、--lambda_c、--lambda_cl | λ_bd、λ_c、λ_cl^j |
Δs、Δc、Δj(α·Δs + β·Δc) | 強度 / 整合性 / ジョイントアダプタ |
eval_gradient_loss.py 内の S、C | 強度 S、勾配整合性 C |
swe_eval_server/ のキュー方式は外部の RL トレーナー(verl、同梱されていない)と統合される。
ワンショットの POST /evaluate エンドポイントは Docker + SWE-bench のみを必要とする。
サーバーは認証なしで 0.0.0.0 にバインドし、モデルが生成したシェルを Docker 内で実行するため、
信頼できるネットワーク上でのみ実行すること。configs/、examples/slurm/、およびそれらの中の outputs/… パスはテンプレートである。
自分の環境に合わせてパスとクラスタディレクティブを編集すること。@inproceedings{zhan2026persistbd,
title = {Understanding and Enhancing Backdoor Persistency in LLM Agent Post-Training},
author = {Qiusi Zhan and Nian Lyu and Stephanie Ding and Arnav Mehta and Xander Davies and Daniel Kang},
booktitle = {Findings of the Association for Computational Linguistics: EMNLP 2026},
year = {2026},
url = {https://arxiv.org/abs/2610.07510}
}