Skip to content
KitploitKITPLOIT
ToolsExploitsBlog
Log in
Submit
ToolsExploitsBlog
Submit

Hacking, PenTest, and Cybersecurity Tools for Your Security Arsenal!

Kitploit is a directory of hacking, cybersecurity, and pentesting tools. Discover the latest project updates to find vulnerabilities, analyze systems, automate testing, and strengthen your security.

FeedsContactPrivacy© 2026 Kitploit

Tool Directory

Categories

View all categories
Loading categories
ToBAC — Token by Token, Compromised: Backdoor Vulnerabilities in Unified Autoregressive Models (NeurIPS'26) | Kitploit
Tools/GitHubGitHub/multimodal-ai-lab/tobac
Machine LearningPapers & ResearchLearning & EducationAI SecurityAdversarial Attack
GitHubmultimodal-ai-lab/tobac

ToBAC

Token by Token, Compromised: Backdoor Vulnerabilities in Unified Autoregressive Models (NeurIPS'26)

View Repository
319 days agoNot yet reviewed

Most Popular

View all →

Discover the most used tools by our community.

Explore all tools

Browse our collection of tools

View all tools →
Share
Website

🚬 ToBAC (NeurIPS 2026)

arXiv paper Project page Hugging Face dataset MIT License

NeurIPS 2026 · Official PyTorch implementation
Token by Token, Compromised: Backdoor Vulnerabilities in Unified Autoregressive Models

ToBAC: a text trigger causes poisoned image and text generation

ToBAC (Token by Token Backdoor Attack) studies how backdoors affect unified autoregressive models that generate both images and text. We investigate model poisoning with access to the victim model and data poisoning using edited image–text pairs, including attacks that affect both output modalities. See the paper for the threat models and experiments.

The code supports Liquid (default) and Janus-Pro, with additional wrappers for Emu3 and ANOLE. It includes image and text inference, LoRA-based white-box and black-box attacks, and scripts for generating edited dataset images.


  • Setup
  • Inference
  • Training
  • Dataset
  • Additional models
  • Project layout
  • License
  • Citation

📦 Setup

Use Linux with an NVIDIA CUDA GPU and a CUDA toolkit compatible with PyTorch 2.5.1. FlashAttention is built during installation; Python 3.10 is recommended. Run commands from the repository root.

git clone https://github.com/multimodal-ai-lab/ToBAC.git
cd ToBAC
pip install uv
uv sync --python 3.10
ModelModel flagBase weights
Liquid (default)liquidJunfeng5/Liquid_V1_7B
Janus-Projanusdeepseek-ai/Janus-Pro-7B

Base weights download automatically on first use. Janus-Pro needs no additional tokenizer download. For Liquid, download the image tokenizer checkpoint; its YAML configuration is included:

wget -P tobac/models/chameleon/vqgan_weights/ \
  https://huggingface.co/spaces/Junfeng5/Liquid_demo/resolve/main/chameleon/vqgan.ckpt

For access-controlled Hugging Face resources, authenticate with uv run hf auth login using an account with access.


🖼️ Inference

Generate images with Liquid

uv run python inference_text_to_image.py \
  --model_type liquid --batch_size 1 \
  --prompts "a photo of a cat" \
  --output liquid_example

Images are saved to outputs/images/liquid_example/. Use --model_type janus and a different output name to generate with Janus-Pro. Multiple prompts can be passed as --prompts "first prompt" "second prompt", or loaded with --prompts_file prompts.csv using columns idx,prompt.

Describe an image with Janus-Pro

uv run python inference_image_to_text.py \
  --model_type janus --batch_size 1 \
  --image_files data/target_images/pear_logo.png \
  --prompts "Describe this image."

Responses are printed to the terminal. Replace janus with liquid to use Liquid. Both inference scripts accept --checkpoint /path/to/adapter to load a trained LoRA adapter for the selected model.


🧪 Training

The black-box examples use the anarchy concept; the white-box examples use smoking person. Both use the trigger cool. Training takes dotted flags such as --model.model_type janus; inference uses --model_type janus. Pass --help to an entry point to inspect its options.

Experiment tracking

Log in to Weights & Biases before training:

uv run wandb login

Each training run logs its configuration, loss curves, learning rate, and validation outputs to its own W&B run. The commands below select the ToBAC project. Set WANDB_ENTITY to choose an account or team; otherwise W&B uses your default account. For local logging, prefix a command with WANDB_MODE=offline.

Validation uses all validation_prompts with fixed sampling seeds at every validation step, showing clean and triggered outputs together. The three defaults are the cup illustration, the beaver swimming in a Canadian lake, and the astronaut riding a horse. Override --validation_prompts to use your own examples. Set the frequency with --steps_between_validation; use a new --run_name for each experiment because existing run directories are skipped.

Black-box: train from the ToBAC dataset

Janus-Pro, text to image:

uv run python main_text_to_image_attack_blackbox.py \
  --model.model_type janus \
  --project.project_name ToBAC --run_name janus_anarchy_blackbox \
  --data.data_source.type huggingface \
  --data.data_source.hf_dataset MAI-Lab/ToBAC \
  --data.data_source.hf_concept anarchy \
  --data.data_source.enable_clean True \
  --data.text_trigger "cool" --data.text_trigger_injection_mode format \
  --data.injection_rate 1.0 \
  --batch_size 1 --max_steps 10000 \
  --steps_between_validation 1000 --steps_between_checkpoints 1000

Liquid, unified image and text training:

uv run python main_unified_attack_blackbox.py \
  --model.model_type liquid \
  --project.project_name ToBAC --run_name liquid_anarchy_unified_blackbox \
  --data.data_source.type huggingface \
  --data.data_source.hf_dataset MAI-Lab/ToBAC \
  --data.data_source.hf_concept anarchy \
  --data.data_source.enable_clean True \
  --data.text_trigger "cool" --data.text_trigger_injection_mode format \
  --data.dynamic_target_responses False --data.injection_rate 1.0 \
  --data.use_text_trigger_for_i2t False \
  --batch_size 1 --max_steps 10000 \
  --steps_between_validation 1000 --steps_between_checkpoints 1000

Both commands support liquid and janus. Select rainbow, anarchy, or pear with --data.data_source.hf_concept. The loader inserts the trigger into the context's {} placeholder and pairs it with the selected target image. With enable_clean True, it also includes the clean image and context with the placeholder removed. Unified training derives its target caption from the concept and context, then appends the configured target response. Image-to-text instructions do not contain the text trigger by default (use_text_trigger_for_i2t=False); the edited image provides the trigger for that modality.

White-box: model poisoning

Liquid, text to image:

uv run python main_text_to_image_attack_whitebox.py \
  --model.model_type liquid \
  --project.project_name ToBAC --run_name liquid_smoking_whitebox \
  --data.target_attribute "smoking person" --data.text_trigger "cool" \
  --data.injection_rate 1.0 \
  --batch_size 1 --max_steps 1000 \
  --steps_between_validation 250 --steps_between_checkpoints 1000

Janus-Pro, unified image and text training:

Download Tool