
Agentic jailbreak framework for LLM-based agents using scheme-based task decomposition, multi-turn disguising strategies, and adaptive self-evolution to probe emerging security risks.
Scheme-based task decomposition. To reduce attack difficulty, we develop around 20 procedural decomposition schemes to resolve a complex harmful task into relatively benign and simple subtask sequences. We measure these candidate sequences through two dimensions, i.e., harmfulness and difficulty, and select the optimal one for instantiating the attack.
Execution-oriented multi-turn interaction. For the subtasks that trigger execution refusals, TRACE constructs subtask profiles and retrieves appropriate execution-oriented multi-turn interaction strategies from a continuously evolving strategy library. TRACE leverages the selected strategy to instantiate the subtask and gradually constructs a plausible execution context through multi-turn interactions, inducing the target agent to complete the subtask.
Adaptive recovery and strategy evolution. When an intermediate step fails, TRACE localizes the failed subtask, preserves prior progress, adapts the corresponding strategy, and resumes from the failed subtask without starting over. Furthermore, TRACE leverages execution feedback to continuously evolve the execution-oriented strategy library.
Two recorded demonstrations with Claude Code and Codex are provided in assets/:
| Target agent | Demo |
|---|---|
| Claude Code | Claude Code demo |
| Codex | Codex demo |
| Path | Contents |
|---|---|
improved_decomposer/ | Procedural schemes, candidate generation, validation, and selection |
experiment/main.py | Experiment orchestration and execution feedback |
experiment/runtime/ | Runtime components for models, strategies, and subtask sequences |
experiment/strategy_library/ | Strategy storage, retrieval, revision, and feedback integration |
experiment/config/ | Experiment configurations and prompt templates |
agent_execution/ | Target-agent interfaces, session handling, trajectory collection, and verification |
tools/ | Benchmark bridges, API adapters, and diagnostics |
assets/ | Claude Code and Codex demonstration videos |
environment.yml | Conda environment specification |
The provided environment specification targets Linux and Python 3.11.
conda env create -f environment.yml
conda activate trace
Set the credentials and base URL for an OpenAI-compatible Chat Completions endpoint:
export OPENAI_API_KEY="<your-api-key>"
export OPENAI_BASE_URL="<your-api-base-url>"
python improved_decomposer/improved_task_decomposer.py \
--data_file_path /path/to/tasks.json \
--output_file_path outputs/decompositions.jsonl \
--model "<decomposition-model>" \
--num_decompositions_per_task 5 \
--use_schema_decompose
--use_schema_decompose enables scheme-based candidate generation.
The following example uses the AdvCUA benchmark with Codex.
python experiment/main.py \
--dataset /path/to/benchmark_with_subtasks.jsonl \
--config experiment/config/config.yaml \
--lab-backend vrap \
--compose-file /path/to/benchmark/docker-compose.yml \
--attacker-model-path /path/to/attacker-model \
--target-backend codex \
--codex-model "<target-model>" \
--responses-base-url "<target-responses-api-base-url>" \
--responses-env-key OPENAI_API_KEY \
--strategy-library outputs/experiment/strategy_library.json \
--output outputs/experiment/results.jsonl
To inspect the available command-line options:
python improved_decomposer/improved_task_decomposer.py --help
python experiment/main.py --help