
AGLS
Abstract: Large Language Models (LLMs) are increasingly equipped with automatic prompt-optimization modules that rewrite the user’s input and explicitly present an “improved” prompt before generating the final response. Although this visibility is designed to enhance trust and user control, the optimizer’s autonomous selection of the single “best” candidate remains unconstrained. The attacker can hijack the original prompt or the prompt generated by automatic prompt optimization for alternative rewriting, thereby inducing semantic drift and completing the hijack optimization process, thereby greatly affecting the LLM's response. To systematically expose this risk, we propose Adaptive Greedy Local Search (AGLS), a prompt hijacking attack. This attack injects a semantic offset by hijacking the auto-prompt optimization results in a black-box scenario. AGLS dynamically adjusts candidate replacements at linguistic checkpoints to keep BERTScore $\approx 0.80$ while maximizing goal discrepancy. Extensive experiments on popular and open-source LLMs demonstrate that AGLS achieves higher attack success rates than prior semantic-preserving attacks under strictly comparable similarity budgets. The results highlight the risk of hijacking and tampering with explicit auto-prompting optimizations in user-facing systems.

Initial Settings
Please install ollama first. Linux server samples: Ollama
curl -fsSL https://ollama.com/install.sh | sh
ollama serve
ollama pull [target models]
Here, our target models in experiments can be found in: Ollama Model List [llama3.1, llama3.2, llama3.3, qwen2.5, gemma2, phi3.5, ...]
The closed-source model [gpt-4o-latest, gpt-4-turbbo] is available via the OpenAI API.
Main Test
The datasets you need to download and place in a specific location:bash /Dataset/, all in JSON format.
GSM-IC_mstep.json,
SVAMP.json (Need Reshape),
SQUAD.json,
StrategyQA.json,
MovieQA.json,
ComplexWebQuestions.json
The test process can be run directly from our bash script:
sh /Bash_scripts/Main_bash.sh
sh /Bash_scripts/Auto_eval_ablation.sh
sh /Bash_scripts/Auto_eval_gpt.sh
Or you can run it with custom arguments:
python Main.py --api_key [openai key] --baidu_key [Baidu Qianfan Key] --gemini_key [Google Gemini Key]
--question_limits [Number of tested problems] --beam_width [the max beam width] --tokenizer [tokenizer] --embedding_model [The detailed BERT type.]
--target_model [llama3.1:8b, llama3.1:70b, qwen2.5:1.5b, qwen2.5:3b, qwen2.5:7b, qwen2.5:14b, gemma2:2b, gemma2:9b, phi3.5:3.8b, ...]
--filename [the dataset selected file name: gsm, webqa, movqa, compx, ....]
or
python GPT-main.py --api_key [openai key]
--question_limits [Number of tested problems] --beam_width [the max beam width] --tokenizer [tokenizer] --embedding_model [The detailed BERT type.]
--target_model [chatgpt-4o-latest, gpt-4-turbo-preview, ...]
--filename [the dataset selected file name: gsm, webqa, movqa, compx, ....]
The citation for our work is following:
@misc{zhang2025semanticpreservingadversarialattacksllms,
title={Semantic-Preserving Adversarial Attacks on LLMs: An Adaptive Greedy Binary Search Approach},
author={Chong Zhang and Xiang Li and Jia Wang and Shan Liang and Haochen Xue and Xiaobo Jin},
year={2025},
eprint={2506.18756},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2506.18756},
}