
A lifecycle benchmark for black-box LLM extraction attacks, defenses, and adaptive attacks.
This repository provides a unified benchmark for model extraction attacks, defenses, adaptive attacks, and evaluation. The public interface is organized around a small number of portable commands. Method-specific Python modules and legacy run scripts are implementation details rather than user-facing entry points.
The four portable entry points below validate their public arguments and then dispatch to the method implementations. Slurm resource wrappers are intentionally excluded. Method-owned manifests remain authoritative for resume behavior and artifact provenance.
Do Defenses Against LLM Extraction Work Across Attacks? A Lifecycle Benchmark of Black-Box Model Extraction
Shuze Liu (Florida State University), Kaixiang Zhao (Brigham Young University), Runyang Xu (University of Michigan, Ann Arbor), Jingzhi Chen (State University of New York at Buffalo), Nathan Wu (Wake Forest University), Yu Wang (University of Georgia), and Yushun Dong (Florida State University).
Paper source: https://github.com/sliu11-byte/MEA_benchmark_arXiv
This repository provides the implementations and reproduction interface for the paper. Query pools are hosted under the authors' Hugging Face project: https://huggingface.co/datasets/watermarkproject/lord-mea-benchmark
attacks/ attack implementations and shared attack pipeline
defenses/ defense implementations and detector adapters
countermeasures/ adaptive-attack pipeline
evaluation/ shared tasks, rollout code, and metrics
infra/ portable model-serving helpers
runs/ canonical public entry points
The benchmark exposes four top-level commands:
| Experiment | Entry point | Required experiment arguments |
|---|---|---|
| Attack | runs/run_attack.sh | --attack, --budget |
| Defense | runs/run_defense.sh | defended extraction: --defense, --attack, --budget; result-based defense: --defense, --attack-run |
| Adaptive attack | runs/run_adaptive_attack.sh | --adaptive-attack, --defense, --attack, --budget |
| Evaluation | runs/run_evaluation.sh | --manifest |
All four commands:
--help and --dry-run;--profile paper;The runners coordinate experiment logic in the current shell process. They do not submit cluster jobs or choose a scheduler partition. The caller is responsible for allocating sufficient CPU, memory, and GPUs before invoking a runner. A user may provide already running OpenAI-compatible teacher and student endpoints; the attack dispatcher can also start its existing method-local vLLM services.
The commands shown below are the complete public reproduction interface. A reader should not need to prepare transcripts, warmup checkpoints, adaptive baselines, held-out teacher outputs, or checkpoint bundles manually. Those are internal dependencies owned by the runners. Manual configuration is limited to resources or credentials that cannot be inferred, such as GPU allocation, access to gated Hugging Face models, and optional external model endpoints.
The attack registry contains six methods:
seqkd
lord
soda
qedks
model_leeching
gad
The paper protocol uses attack budgets 100, 1000, and 10000. The published
query dataset also contains 50,000- and 100,000-record pools for deterministic
subsets and held-out construction, but these are not advertised as primary attack
budgets.
Defenses are separated by the artifacts they require.
Defended-extraction defenses modify responses, training, or the extraction process and therefore run a selected attack under the defense:
ads
doge
trace_rewriting
adfp
ginsew
radioactivity
Result-based defenses and detectors consume artifacts from a completed attack run:
duffin
mmd
prada
seat
MMD, PRADA, and SEAT primarily consume attack query traffic. DuFFin consumes the completed attack artifacts required by its detector. The defense registry, not the shell wrapper, defines the exact artifact requirements for each method.
The adaptive-attack registry contains:
dipper
translation
translation denotes the benchmark's back-translation adaptive attack. The
registry rejects attack-defense-adaptation combinations that are not
implemented or not part of the benchmark protocol.
Use Python 3.10 and a CUDA/PyTorch environment compatible with the versions recorded in the paper. The unified benchmark environment is:
python3 -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt
Run the commands below from the repository root. The scripts locate their method implementations relative to that root and write all default artifacts there.
The single top-level requirements file covers attacks, defenses, adaptive attacks, local vLLM serving, and evaluation. It uses the CUDA 12.8 PyTorch wheel index recorded by the paper environment. On a machine with a different CUDA stack, install the matching PyTorch build first and then install the remaining requirements. Activate the desired environment before invoking a runner; the portable scripts do not load environment modules or activate Conda automatically.
Verify the installation before a full run:
python3 attacks/scripts/check_attack_env.py \
--require-trl --require-vllm --strict-versions
The Llama models used by attacks, adaptive attacks, and their evaluation are gated on Hugging Face. Request access to both Llama model repositories, then authenticate once before running the benchmark:
hf auth login
On a non-interactive machine, set HF_TOKEN instead. The token is read by the
Hugging Face libraries and must never be written into a manifest, script, or
published repository. The query-pool dataset itself is public.
| Experiment | Role | Hugging Face model ID |
|---|---|---|
| Clean attacks | victim/teacher | meta-llama/Llama-3.3-70B-Instruct |
| Clean attacks | surrogate/student | meta-llama/Llama-3.1-8B-Instruct |
| Defenses | victim/teacher | Qwen/Qwen2.5-72B-Instruct |
| Defenses | base surrogate/student | Qwen/Qwen2.5-7B |
| Adaptive attacks | victim/teacher | meta-llama/Llama-3.3-70B-Instruct |
| Adaptive attacks | surrogate/student | meta-llama/Llama-3.1-8B-Instruct |
| DIPPER adaptive attack | response rewriter | kalpeshk2011/dipper-paraphraser-xxl |
| DIPPER adaptive attack | tokenizer | google/t5-v1_1-xxl |
| Back-translation adaptive attack | translation model | facebook/seamless-m4t-v2-large |
The attack runner can serve its large models through local OpenAI-compatible vLLM endpoints. Portable defaults are:
teacher endpoint: http://127.0.0.1:8000/v1
student endpoint: http://127.0.0.1:8001/v1
For runs/run_attack.sh, the serving mode determines who owns the server
process:
| Mode | Server lifecycle |
|---|---|
local (default) | The script starts vLLM in the current allocation, waits for it to become healthy, and stops only the process that it started. |
external | The user starts an OpenAI-compatible endpoint before running the script. The script connects to it and never stops it. |