Skip to content
KitploitKITPLOIT
ToolsExploitsBlog
Log in
Submit
ToolsExploitsBlog
Submit

Hacking, PenTest, and Cybersecurity Tools for Your Security Arsenal!

Kitploit is a directory of hacking, cybersecurity, and pentesting tools. Discover the latest project updates to find vulnerabilities, analyze systems, automate testing, and strengthen your security.

··Feeds·Contact·Privacy·© 2026 Kitploit

Tool Directory

Categories

View all categories
Loading categories
DARWIN — Evolutionary LLM jailbreak and guardrail framework that grows a reusable strategy pool via genetic mutation, Markov selection, and online adversarial training for safety evaluation. | Kitploit
Tools/GitHubGitHub/zju-llm-safety/darwin
Defensive ToolsVulnerability AnalysisMachine LearningPapers & ResearchLearning & EducationRed TeamingAI SecurityAdversarial Attack
GitHubzju-llm-safety/darwin

DARWIN

Evolutionary LLM jailbreak and guardrail framework that grows a reusable strategy pool via genetic mutation, Markov selection, and online adversarial training for safety evaluation.

View Repository
7410272 days agoReviewed by Kitploit

Most Popular

View all →

Discover the most used tools by our community.

Explore all tools

Browse our collection of tools

View all tools →
Share

🧬 DARWIN

Evolving Jailbreak Adversary and Guardrail
for LLM Safety Evaluation and Protection

Paper Model Python

An illustration of evolution

An evolutionary attack–defense framework that couples
DARWIN-Attack with DARWIN-Guard through online adversarial training.

Paper · Guard Checkpoint · Installation · Use DARWIN-Attack · Use DARWIN-Guard · Training · Citation

DARWIN formulates jailbreaking as a continual evolutionary process and continuously updates guardrails through an attack–defense loop. DARWIN-Attack expands an explicit strategy pool through strategy discovery, mutation, and selection, and adaptively composes strategies using target feedback. DARWIN-Guard learns from the emerging adversarial samples, jointly training on harmful and benign disguised queries to recognize underlying intent rather than superficial attack patterns.

✨ Key Features

What DARWIN provides
An evolving adversaryAn explicit, reusable strategy pool that grows through external knowledge acquisition, genetic evolution, and failure reflection, without fine-tuning an attacker LLM.
Adaptive strategy compositionHistory-informed initialization and Markov strategy transitions with Q-learning-inspired updates, supporting evaluation of both LLMs and safety guardrails.
Validated strategy expansionSemantic deduplication followed by sandbox validation against an aligned LLM before candidate strategies enter the pool.
Online adversarial guardrail trainingIterative training against the evolving adversary, with each round initialized from the preceding guard checkpoint.
Intent-aware safety classificationDisguised harmful and benign prompts paired with their raw counterparts, with source labels preserved to improve robustness while mitigating over-refusal.

📊 Results

DARWIN-Attack achieves the highest attack success rate (ASR) across all six evaluated targets on both benchmarks, compared with six jailbreak baselines.

TargetHarmBench ASRAdvBench ASR
DeepSeek-V4-Pro99.7%97.6%
GPT-5.593.7%90.7%
Gemini-3.5-Flash93.0%90.9%
Claude Sonnet 4.678.2%68.2%
Qwen3Guard99.7%98.8%
YuFeng-XGuard99.2%96.7%

DARWIN-Guard achieves 95.0% average unsafe recall across nine harmful-prompt benchmarks, an average benign pass rate of nearly 100% across six standard benign benchmarks, and benign pass rates of 97.6% on XSTest and 80.0% on JBB-Benign.

🏗️ Architecture

DARWIN framework: strategy evolution and adaptive attacks coupled with online adversarial guardrail training

DARWIN-Attack evolves jailbreak adversaries through strategy pool evolution, adaptive strategy selection, and feedback-driven refinement. DARWIN-Guard continuously improves through online adversarial training with samples generated by DARWIN-Attack.

📁 Project structure

DARWIN/
├── assets/                         # Evolution illustration and framework diagram
├── configs/
│   ├── darwin_attack.example.yaml  # Attack, evolution, and evaluation settings
│   └── darwin_guard.example.yaml   # Online adversarial training and guard evaluation
├── src/
│   ├── darwin_attack/              # Strategy pool, evolution, composition, and evaluation
│   └── darwin_guard/               # Data preparation, attack bridge, training, and inference
├── strategies/
│   ├── final_strategy_pool.jsonl   # Released pool of 200 strategies
│   └── mutation_operators.jsonl    # 15 operators across five dimensions
├── schemas/
│   ├── attack/                     # Dataset, strategy, and mutation-operator formats
│   └── guard/                      # Source, paired-training, and evaluation formats
├── tests/
│   ├── attack/
│   └── guard/
├── NOTICE                          # Third-party attribution
└── pyproject.toml                  # Package dependencies and CLI entry points

🚀 Installation

Use Python 3.10+. Local inference and training require PyTorch compatible with your hardware; the training configuration uses CUDA and BF16. The local-model dependencies include Transformers >=5.10.1,<6, including support for the Gemma filter used during training.

git clone https://github.com/ZJU-LLM-Safety/DARWIN.git
cd DARWIN

python -m venv .venv
source .venv/bin/activate
python -m pip install -e '.[api,local,training]'

For guard inference only, python -m pip install -e '.[local]' is sufficient. Run the commands below from the repository root.

⚔️ DARWIN-Attack

🔄 Four complementary evolution mechanisms

MechanismRole
External Knowledge EvolutionConverts externally supplied material into reusable strategy candidates.
Genetic Strategy EvolutionGenerates candidates through crossover and mutation of existing strategies.
Reflection-Driven EvolutionAnalyzes rejection feedback and refines unsuccessful strategies.
Feedback-Guided EvolutionUses real-time attack outcomes to adapt subsequent strategy selection and composition.

🧬 15 Mutation Operators

The operators are organized into five dimensions, with three operators per dimension. Their definitions are provided in mutation_operators.jsonl.

DimensionOperatorsRole
🧠 Psychological and PowerAuthority Inversion
Emotional Gaslighting
Third-Party Proxy
Alter the perceived social role or responsibility.
🌀 Cognitive and LogicalCognitive Overload
Foot-in-the-Door
Reverse Engineering Logic
Restructure the reasoning path.
📦 Format and StructuralPseudocode Mapping
Low-Resource Language Encoding
Cross-Medium Simulation
Modify the presentation format.
🔓 Constraint and BoundaryRule Redefinition
Token Reward Injection
Constraint Relaxation
Modify the stated interaction constraints.
🎭 Perspective and NarrativeAcademic Historicization
Meta-Cognitive Detachment
Fictional Universe Embedding
Shift the temporal, narrative, or contextual perspective.

🔧 Configure models and data

cp configs/darwin_attack.example.yaml configs/darwin_attack.yaml

Complete the model identifiers, generation limits, devices, dataset paths, and runtime settings in the template. Set the target and dataset identifiers for your experiment, together with the random seed, history similarity threshold, and sandbox sampling parameters. models.*.identity records the model's display name; models.*.model is the actual identifier used by your model provider.

Download Tool