Skip to content
KitploitKITPLOIT
ToolsExploitsBlog
Log in
Submit
ToolsExploitsBlog
Submit

Hacking, PenTest, and Cybersecurity Tools for Your Security Arsenal!

Kitploit is a directory of hacking, cybersecurity, and pentesting tools. Discover the latest project updates to find vulnerabilities, analyze systems, automate testing, and strengthen your security.

FeedsContactPrivacy© 2026 Kitploit

Tool Directory

Categories

View all categories
Loading categories
MEA-Bench — A lifecycle benchmark for black-box LLM extraction attacks, defenses, and adaptive attacks. | Kitploit
Tools/GitHubGitHub/sliu11-byte/mea-bench
Vulnerability AnalysisMachine LearningPapers & ResearchLearning & EducationAI SecurityAdversarial Attack
GitHubsliu11-byte/mea-bench

MEA-Bench

A lifecycle benchmark for black-box LLM extraction attacks, defenses, and adaptive attacks.

View Repository
41 day agoNot yet reviewed

Most Popular

View all →

Discover the most used tools by our community.

Explore all tools

Browse our collection of tools

View all tools →
Share

MEA-Bench: A Benchmark for Model Extraction Attacks

This repository provides a unified benchmark for model extraction attacks, defenses, adaptive attacks, and evaluation. The public interface is organized around a small number of portable commands. Method-specific Python modules and legacy run scripts are implementation details rather than user-facing entry points.

The four portable entry points below validate their public arguments and then dispatch to the method implementations. Slurm resource wrappers are intentionally excluded. Method-owned manifests remain authoritative for resume behavior and artifact provenance.

Paper and Authors

Do Defenses Against LLM Extraction Work Across Attacks? A Lifecycle Benchmark of Black-Box Model Extraction

Shuze Liu (Florida State University), Kaixiang Zhao (Brigham Young University), Runyang Xu (University of Michigan, Ann Arbor), Jingzhi Chen (State University of New York at Buffalo), Nathan Wu (Wake Forest University), Yu Wang (University of Georgia), and Yushun Dong (Florida State University).

Paper source: https://github.com/sliu11-byte/MEA_benchmark_arXiv

This repository provides the implementations and reproduction interface for the paper. Query pools are hosted under the authors' Hugging Face project: https://huggingface.co/datasets/watermarkproject/lord-mea-benchmark

Repository Layout

attacks/          attack implementations and shared attack pipeline
defenses/         defense implementations and detector adapters
countermeasures/  adaptive-attack pipeline
evaluation/       shared tasks, rollout code, and metrics
infra/            portable model-serving helpers
runs/             canonical public entry points

Canonical Public Interface

The benchmark exposes four top-level commands:

ExperimentEntry pointRequired experiment arguments
Attackruns/run_attack.sh--attack, --budget
Defenseruns/run_defense.shdefended extraction: --defense, --attack, --budget; result-based defense: --defense, --attack-run
Adaptive attackruns/run_adaptive_attack.sh--adaptive-attack, --defense, --attack, --budget
Evaluationruns/run_evaluation.sh--manifest

All four commands:

  • work as ordinary shell commands without Slurm;
  • accept --help and --dry-run;
  • validate the requested experiment before starting;
  • automatically create, discover, validate, and reuse all prerequisite artifacts;
  • use deterministic benchmark defaults under --profile paper;
  • preserve each method's existing resume and artifact-reuse behavior;
  • write or preserve machine-readable run metadata and input provenance;
  • avoid embedded usernames, cluster accounts, email addresses, or site-specific paths.

The runners coordinate experiment logic in the current shell process. They do not submit cluster jobs or choose a scheduler partition. The caller is responsible for allocating sufficient CPU, memory, and GPUs before invoking a runner. A user may provide already running OpenAI-compatible teacher and student endpoints; the attack dispatcher can also start its existing method-local vLLM services.

The commands shown below are the complete public reproduction interface. A reader should not need to prepare transcripts, warmup checkpoints, adaptive baselines, held-out teacher outputs, or checkpoint bundles manually. Those are internal dependencies owned by the runners. Manual configuration is limited to resources or credentials that cannot be inferred, such as GPU allocation, access to gated Hugging Face models, and optional external model endpoints.

Supported Methods

Attacks

The attack registry contains six methods:

seqkd
lord
soda
qedks
model_leeching
gad

The paper protocol uses attack budgets 100, 1000, and 10000. The published query dataset also contains 50,000- and 100,000-record pools for deterministic subsets and held-out construction, but these are not advertised as primary attack budgets.

Defenses

Defenses are separated by the artifacts they require.

Defended-extraction defenses modify responses, training, or the extraction process and therefore run a selected attack under the defense:

ads
doge
trace_rewriting
adfp
ginsew
radioactivity

Result-based defenses and detectors consume artifacts from a completed attack run:

duffin
mmd
prada
seat

MMD, PRADA, and SEAT primarily consume attack query traffic. DuFFin consumes the completed attack artifacts required by its detector. The defense registry, not the shell wrapper, defines the exact artifact requirements for each method.

Adaptive Attacks

The adaptive-attack registry contains:

dipper
translation

translation denotes the benchmark's back-translation adaptive attack. The registry rejects attack-defense-adaptation combinations that are not implemented or not part of the benchmark protocol.

Setup

Use Python 3.10 and a CUDA/PyTorch environment compatible with the versions recorded in the paper. The unified benchmark environment is:

python3 -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt

Run the commands below from the repository root. The scripts locate their method implementations relative to that root and write all default artifacts there.

The single top-level requirements file covers attacks, defenses, adaptive attacks, local vLLM serving, and evaluation. It uses the CUDA 12.8 PyTorch wheel index recorded by the paper environment. On a machine with a different CUDA stack, install the matching PyTorch build first and then install the remaining requirements. Activate the desired environment before invoking a runner; the portable scripts do not load environment modules or activate Conda automatically.

Verify the installation before a full run:

python3 attacks/scripts/check_attack_env.py \
  --require-trl --require-vllm --strict-versions

The Llama models used by attacks, adaptive attacks, and their evaluation are gated on Hugging Face. Request access to both Llama model repositories, then authenticate once before running the benchmark:

hf auth login

On a non-interactive machine, set HF_TOKEN instead. The token is read by the Hugging Face libraries and must never be written into a manifest, script, or published repository. The query-pool dataset itself is public.

Models Used in the Paper

ExperimentRoleHugging Face model ID
Clean attacksvictim/teachermeta-llama/Llama-3.3-70B-Instruct
Clean attackssurrogate/studentmeta-llama/Llama-3.1-8B-Instruct
Defensesvictim/teacherQwen/Qwen2.5-72B-Instruct
Defensesbase surrogate/studentQwen/Qwen2.5-7B
Adaptive attacksvictim/teachermeta-llama/Llama-3.3-70B-Instruct
Adaptive attackssurrogate/studentmeta-llama/Llama-3.1-8B-Instruct
DIPPER adaptive attackresponse rewriterkalpeshk2011/dipper-paraphraser-xxl
DIPPER adaptive attacktokenizergoogle/t5-v1_1-xxl
Back-translation adaptive attacktranslation modelfacebook/seamless-m4t-v2-large

The attack runner can serve its large models through local OpenAI-compatible vLLM endpoints. Portable defaults are:

teacher endpoint: http://127.0.0.1:8000/v1
student endpoint: http://127.0.0.1:8001/v1

For runs/run_attack.sh, the serving mode determines who owns the server process:

ModeServer lifecycle
local (default)The script starts vLLM in the current allocation, waits for it to become healthy, and stops only the process that it started.
externalThe user starts an OpenAI-compatible endpoint before running the script. The script connects to it and never stops it.
Download Tool