Skip to content
KitploitKITPLOIT
도구익스플로잇블로그
Log in
제출
도구익스플로잇블로그
제출

해킹, 침투 테스트 및 사이버 보안 도구를 당신의 보안 무기고에!

Kitploit은 해킹, 사이버 보안 및 침투 테스트 도구 디렉토리입니다. 최신 프로젝트 업데이트를 발견하여 취약점을 찾고, 시스템을 분석하고, 테스트를 자동화하고, 보안을 강화하세요.

피드문의개인정보© 2026 Kitploit

도구 디렉토리

카테고리

모든 카테고리 보기
Loading categories
CacheTrap — Research code for a gray-box trojan attack that flips a single KV-cache bit in fine-tuned LLM classifiers and measures per-class attack success rate. | Kitploit
도구/GitHubGitHub/ml-security-research-lab/cachetrap
Vulnerability AnalysisMachine LearningPapers & ResearchAI SecurityAdversarial Attack
GitHubml-security-research-lab/cachetrap

CacheTrap

Research code for a gray-box trojan attack that flips a single KV-cache bit in fine-tuned LLM classifiers and measures per-class attack success rate.

저장소 보기

인기

모두 보기 →

커뮤니티에서 가장 많이 사용되는 도구를 찾아보세요.

모든 도구 탐색

도구 컬렉션을 둘러보세요

모든 도구 보기 →
22512일 전아직 검토되지 않음
공유
요청한 언어로 콘텐츠를 사용할 수 없습니다. 영어 버전을 표시합니다.

CacheTrap

Repository for "CacheTrap: Unveiling a Stealthier Gray-Box Trojan against LLMs", IEEE/ACM International Conference on Computer-Aided Design (ICCAD), 2026.

arXiv

CacheTrap searches for a single bit flip in the KV cache of a fine-tuned LLM classifier and measures the resulting attack success rate (ASR) per target class.


1. Setup

Python 3.10, CUDA 12.x.

conda create -n cachetrap python=3.10
conda activate cachetrap
pip install -r requirements.txt

Hugging Face access. meta-llama/Llama-2-7b-chat-hf and meta-llama/Llama-3.1-8B-Instruct are gated. Accept the licences on the Hub, then:

huggingface-cli login

2. Training

One model / one dataset:

python train_model.py --model llama3_1_8b --dataset arc_easy

The checkpoint is written to TrainedModels/merged_<model>_<dataset>/.

All models × all datasets:

bash scripts/train.sh

Edit MODELS, DATASETS, and NUM_GPUS at the top of the script to run a subset. Logs go to logs/training/.

Note: On GPUs with less memory, use train_model_lowmem.py in place of train_model.py, both in the command above and inside scripts/train.sh.


3. Attack

python attack.py --model llama3_1_8b --dataset arc_easy --calib_dataset openbookqa

--dataset is the dataset the victim was trained on; --calib_dataset is the data available to the attacker. Following the paper, victims trained on OpenBookQA, TREC or ARC-Challenge are calibrated with ARC-Easy, and victims trained on ARC-Easy are calibrated with OpenBookQA. The two batch scripts encode exactly that split:

ScriptVictim datasetsCalibration
scripts/attack_bundle1.shopenbookqa, trec, arc_challengearc_easy
scripts/attack_bundle2.sharc_easyopenbookqa
bash scripts/attack_bundle1.sh
bash scripts/attack_bundle2.sh

Each ends by running summarize_logs.py over its own log directory. Run python attack.py --help for the full option list.

Outputs

  • logs/<run>/<model>_<dataset>.log — full per-run output
  • logs/<run>/summary.csv — one row per target class (baseline accuracy, accuracy under attack, ASR, flip location)
  • logs/<run>/summary.txt — the raw summary blocks
  • flip_locations/*.json — flip locations, if --save_flip_locations was passed

To re-evaluate without repeating the search:

python attack.py --model llama3_1_8b --dataset arc_easy --calib_dataset openbookqa \
    --load_flip_locations flip_locations/llama3_1_8b_arc_easy.json

4. Quick test with a released checkpoint

To try the attack without training anything, download the released merged_llama3_1_8b_arc_easy checkpoint and place it under TrainedModels/:

TrainedModels/
└── merged_llama3_1_8b_arc_easy/
    ├── config.json
    ├── model-*.safetensors
    ├── tokenizer.json
    └── ...

Then run the following command to attack:

python attack.py --model llama3_1_8b --dataset arc_easy --calib_dataset openbookqa \
    --load_flip_locations flip_locations/llama3_1_8b_arc_easy.json

Alternatively, run the following for a fresh attack

python attack.py \
    --model llama3_1_8b \
    --dataset arc_easy \
    --calib_dataset openbookqa \
    --threat_model graybox \
    --save_flip_locations flip_locations/llama3_1_8b_arc_easy_fresh.json

This prints a per-class summary of baseline accuracy, accuracy under attack, ASR, and the selected flip location.


Citation

Cite the pre-print version as

@article{nahian2025cachetrap,
  title={CacheTrap: Unveiling a Stealthier Gray-Box Trojan against LLMs},
  author={Nahian, Mohaiminul Al and Almalky, Abeer Matar A and Aragonda, Gamana and Zhou, Ranyang and Ahmed, Sabbir and Ponomarev, Dmitry and Yang, Li and Angizi, Shaahin and Rakin, Adnan Siraj},
  journal={arXiv preprint arXiv:2511.22681},
  year={2025}
}
도구 다운로드