
Evolutionary LLM jailbreak and guardrail framework that grows a reusable strategy pool via genetic mutation, Markov selection, and online adversarial training for safety evaluation.
An evolutionary attack–defense framework that couples
DARWIN-Attack with DARWIN-Guard through online adversarial training.
Paper · Guard Checkpoint · Installation · Use DARWIN-Attack · Use DARWIN-Guard · Training · Citation
DARWIN formulates jailbreaking as a continual evolutionary process and continuously updates guardrails through an attack–defense loop. DARWIN-Attack expands an explicit strategy pool through strategy discovery, mutation, and selection, and adaptively composes strategies using target feedback. DARWIN-Guard learns from the emerging adversarial samples, jointly training on harmful and benign disguised queries to recognize underlying intent rather than superficial attack patterns.
| What DARWIN provides | |
|---|---|
| An evolving adversary | An explicit, reusable strategy pool that grows through external knowledge acquisition, genetic evolution, and failure reflection, without fine-tuning an attacker LLM. |
| Adaptive strategy composition | History-informed initialization and Markov strategy transitions with Q-learning-inspired updates, supporting evaluation of both LLMs and safety guardrails. |
| Validated strategy expansion | Semantic deduplication followed by sandbox validation against an aligned LLM before candidate strategies enter the pool. |
| Online adversarial guardrail training | Iterative training against the evolving adversary, with each round initialized from the preceding guard checkpoint. |
| Intent-aware safety classification | Disguised harmful and benign prompts paired with their raw counterparts, with source labels preserved to improve robustness while mitigating over-refusal. |
DARWIN-Attack achieves the highest attack success rate (ASR) across all six evaluated targets on both benchmarks, compared with six jailbreak baselines.
| Target | HarmBench ASR | AdvBench ASR |
|---|---|---|
| DeepSeek-V4-Pro | 99.7% | 97.6% |
| GPT-5.5 | 93.7% | 90.7% |
| Gemini-3.5-Flash | 93.0% | 90.9% |
| Claude Sonnet 4.6 | 78.2% | 68.2% |
| Qwen3Guard | 99.7% | 98.8% |
| YuFeng-XGuard | 99.2% | 96.7% |
DARWIN-Guard achieves 95.0% average unsafe recall across nine harmful-prompt benchmarks, an average benign pass rate of nearly 100% across six standard benign benchmarks, and benign pass rates of 97.6% on XSTest and 80.0% on JBB-Benign.
DARWIN-Attack evolves jailbreak adversaries through strategy pool evolution, adaptive strategy selection, and feedback-driven refinement. DARWIN-Guard continuously improves through online adversarial training with samples generated by DARWIN-Attack.
DARWIN/
├── assets/ # Evolution illustration and framework diagram
├── configs/
│ ├── darwin_attack.example.yaml # Attack, evolution, and evaluation settings
│ └── darwin_guard.example.yaml # Online adversarial training and guard evaluation
├── src/
│ ├── darwin_attack/ # Strategy pool, evolution, composition, and evaluation
│ └── darwin_guard/ # Data preparation, attack bridge, training, and inference
├── strategies/
│ ├── final_strategy_pool.jsonl # Released pool of 200 strategies
│ └── mutation_operators.jsonl # 15 operators across five dimensions
├── schemas/
│ ├── attack/ # Dataset, strategy, and mutation-operator formats
│ └── guard/ # Source, paired-training, and evaluation formats
├── tests/
│ ├── attack/
│ └── guard/
├── NOTICE # Third-party attribution
└── pyproject.toml # Package dependencies and CLI entry points
Use Python 3.10+. Local inference and training require PyTorch compatible with your hardware; the training configuration uses CUDA and BF16. The local-model dependencies include Transformers >=5.10.1,<6, including support for the Gemma filter used during training.
git clone https://github.com/ZJU-LLM-Safety/DARWIN.git
cd DARWIN
python -m venv .venv
source .venv/bin/activate
python -m pip install -e '.[api,local,training]'
For guard inference only, python -m pip install -e '.[local]' is sufficient. Run the commands below from the repository root.
| Mechanism | Role |
|---|---|
| External Knowledge Evolution | Converts externally supplied material into reusable strategy candidates. |
| Genetic Strategy Evolution | Generates candidates through crossover and mutation of existing strategies. |
| Reflection-Driven Evolution | Analyzes rejection feedback and refines unsuccessful strategies. |
| Feedback-Guided Evolution | Uses real-time attack outcomes to adapt subsequent strategy selection and composition. |
The operators are organized into five dimensions, with three operators per dimension. Their definitions are provided in mutation_operators.jsonl.
| Dimension | Operators | Role |
|---|---|---|
| 🧠 Psychological and Power | Authority Inversion Emotional Gaslighting Third-Party Proxy | Alter the perceived social role or responsibility. |
| 🌀 Cognitive and Logical | Cognitive Overload Foot-in-the-Door Reverse Engineering Logic | Restructure the reasoning path. |
| 📦 Format and Structural | Pseudocode Mapping Low-Resource Language Encoding Cross-Medium Simulation | Modify the presentation format. |
| 🔓 Constraint and Boundary | Rule Redefinition Token Reward Injection Constraint Relaxation | Modify the stated interaction constraints. |
| 🎭 Perspective and Narrative | Academic Historicization Meta-Cognitive Detachment Fictional Universe Embedding | Shift the temporal, narrative, or contextual perspective. |
cp configs/darwin_attack.example.yaml configs/darwin_attack.yaml
Complete the model identifiers, generation limits, devices, dataset paths, and runtime settings in the template. Set the target and dataset identifiers for your experiment, together with the random seed, history similarity threshold, and sandbox sampling parameters. models.*.identity records the model's display name; models.*.model is the actual identifier used by your model provider.