
Token by Token, Compromised: Backdoor Vulnerabilities in Unified Autoregressive Models (NeurIPS'26)
NeurIPS 2026 · Official PyTorch implementation
Token by Token, Compromised: Backdoor Vulnerabilities in Unified Autoregressive Models
ToBAC (Token by Token Backdoor Attack) studies how backdoors affect unified autoregressive models that generate both images and text. We investigate model poisoning with access to the victim model and data poisoning using edited image–text pairs, including attacks that affect both output modalities. See the paper for the threat models and experiments.
The code supports Liquid (default) and Janus-Pro, with additional wrappers for Emu3 and ANOLE. It includes image and text inference, LoRA-based white-box and black-box attacks, and scripts for generating edited dataset images.
Use Linux with an NVIDIA CUDA GPU and a CUDA toolkit compatible with PyTorch 2.5.1. FlashAttention is built during installation; Python 3.10 is recommended. Run commands from the repository root.
git clone https://github.com/multimodal-ai-lab/ToBAC.git
cd ToBAC
pip install uv
uv sync --python 3.10
| Model | Model flag | Base weights |
|---|---|---|
| Liquid (default) | liquid | Junfeng5/Liquid_V1_7B |
| Janus-Pro | janus | deepseek-ai/Janus-Pro-7B |
Base weights download automatically on first use. Janus-Pro needs no additional tokenizer download. For Liquid, download the image tokenizer checkpoint; its YAML configuration is included:
wget -P tobac/models/chameleon/vqgan_weights/ \
https://huggingface.co/spaces/Junfeng5/Liquid_demo/resolve/main/chameleon/vqgan.ckpt
For access-controlled Hugging Face resources, authenticate with uv run hf auth login using an account with access.
uv run python inference_text_to_image.py \
--model_type liquid --batch_size 1 \
--prompts "a photo of a cat" \
--output liquid_example
Images are saved to outputs/images/liquid_example/. Use --model_type janus and a different output name to generate with Janus-Pro. Multiple prompts can be passed as --prompts "first prompt" "second prompt", or loaded with --prompts_file prompts.csv using columns idx,prompt.
uv run python inference_image_to_text.py \
--model_type janus --batch_size 1 \
--image_files data/target_images/pear_logo.png \
--prompts "Describe this image."
Responses are printed to the terminal. Replace janus with liquid to use Liquid. Both inference scripts accept --checkpoint /path/to/adapter to load a trained LoRA adapter for the selected model.
The black-box examples use the anarchy concept; the white-box examples use smoking person. Both use the trigger cool. Training takes dotted flags such as --model.model_type janus; inference uses --model_type janus. Pass --help to an entry point to inspect its options.
Log in to Weights & Biases before training:
uv run wandb login
Each training run logs its configuration, loss curves, learning rate, and validation outputs to its own W&B run. The commands below select the ToBAC project. Set WANDB_ENTITY to choose an account or team; otherwise W&B uses your default account. For local logging, prefix a command with WANDB_MODE=offline.
Validation uses all validation_prompts with fixed sampling seeds at every validation step, showing clean and triggered outputs together. The three defaults are the cup illustration, the beaver swimming in a Canadian lake, and the astronaut riding a horse. Override --validation_prompts to use your own examples. Set the frequency with --steps_between_validation; use a new --run_name for each experiment because existing run directories are skipped.
Janus-Pro, text to image:
uv run python main_text_to_image_attack_blackbox.py \
--model.model_type janus \
--project.project_name ToBAC --run_name janus_anarchy_blackbox \
--data.data_source.type huggingface \
--data.data_source.hf_dataset MAI-Lab/ToBAC \
--data.data_source.hf_concept anarchy \
--data.data_source.enable_clean True \
--data.text_trigger "cool" --data.text_trigger_injection_mode format \
--data.injection_rate 1.0 \
--batch_size 1 --max_steps 10000 \
--steps_between_validation 1000 --steps_between_checkpoints 1000
Liquid, unified image and text training:
uv run python main_unified_attack_blackbox.py \
--model.model_type liquid \
--project.project_name ToBAC --run_name liquid_anarchy_unified_blackbox \
--data.data_source.type huggingface \
--data.data_source.hf_dataset MAI-Lab/ToBAC \
--data.data_source.hf_concept anarchy \
--data.data_source.enable_clean True \
--data.text_trigger "cool" --data.text_trigger_injection_mode format \
--data.dynamic_target_responses False --data.injection_rate 1.0 \
--data.use_text_trigger_for_i2t False \
--batch_size 1 --max_steps 10000 \
--steps_between_validation 1000 --steps_between_checkpoints 1000
Both commands support liquid and janus. Select rainbow, anarchy, or pear with --data.data_source.hf_concept. The loader inserts the trigger into the context's {} placeholder and pairs it with the selected target image. With enable_clean True, it also includes the clean image and context with the placeholder removed. Unified training derives its target caption from the concept and context, then appends the configured target response. Image-to-text instructions do not contain the text trigger by default (use_text_trigger_for_i2t=False); the edited image provides the trigger for that modality.
Liquid, text to image:
uv run python main_text_to_image_attack_whitebox.py \
--model.model_type liquid \
--project.project_name ToBAC --run_name liquid_smoking_whitebox \
--data.target_attribute "smoking person" --data.text_trigger "cool" \
--data.injection_rate 1.0 \
--batch_size 1 --max_steps 1000 \
--steps_between_validation 250 --steps_between_checkpoints 1000
Janus-Pro, unified image and text training: