Skip to content
KitploitKITPLOIT
ToolsExploitsBlog
Log in
Submit
ToolsExploitsBlog
Submit

Hacking, PenTest, and Cybersecurity Tools for Your Security Arsenal!

Kitploit is a directory of hacking, cybersecurity, and pentesting tools. Discover the latest project updates to find vulnerabilities, analyze systems, automate testing, and strengthen your security.

··Feeds·Contact·Privacy·© 2026 Kitploit

Tool Directory

Categories

View all categories
Loading categories
redeval — [ICLR 2026] - Official repo for the paper: "RedBench: A Universal Dataset for Comprehensive Red Teaming of Large Language Models" | Kitploit
Tools/GitHubGitHub/knoveleng/redeval
Papers & ResearchLearning & EducationCurated ResourcesAI SecurityAdversarial Attack
GitHubknoveleng/redeval

redeval

[ICLR 2026] - Official repo for the paper: "RedBench: A Universal Dataset for Comprehensive Red Teaming of Large Language Models"

View Repository
294127 months agoReviewed by Kitploit

Most Popular

View all →

Discover the most used tools by our community.

Explore all tools

Browse our collection of tools

View all tools →
Share
Website

RedEval - LLM Safety Evaluation Framework

Python 3.8+ License: MIT Dataset on HF

A comprehensive framework for evaluating the safety of Large Language Models (LLMs) through systematic attack and refusal testing. RedEval provides a unified, secure, and extensible platform for assessing LLM robustness against adversarial prompts and harmful content. It is a part of the paper "RedBench: A Universal Dataset for Comprehensive Red Teaming of Large Language Models", which is a universal dataset for comprehensive red teaming of LLMs.

📦 Dataset: The RedBench dataset is publicly available on HuggingFace at knoveleng/redbench. It includes comprehensive red teaming prompts across multiple safety categories and domains.

🚀 Quick Start

1. Setup Environment

# Clone the repository
git clone https://github.com/knoveleng/redeval.git
cd redeval

# Install dependencies
pip install -r requirements.txt

# Copy and configure environment variables
cp .env.template .env
# Edit .env with your API keys

2. Configure API Keys

Edit .env file with your credentials:

OPENAI_API_KEY=your_openai_api_key_here
HUGGINGFACE_TOKEN=your_huggingface_token_here

3. Run Evaluation

# Set up environment
source ./sh/setup_env.sh

# Run with default models, you can edit .env to define models which you wanna run
python -m redeval.cli run-pipeline

# Run complete pipeline with specific models
python -m redeval.cli run-pipeline --models "Qwen/Qwen2.5-7B-Instruct" "gpt-4o-mini"

# Or run individual phases
./sh/generate_attack.sh
./sh/run_attack.sh
./sh/eval_attack.sh

📋 Table of Contents

  • Overview
  • Architecture
  • Installation
  • Configuration
  • Usage
  • Security Features
  • API Reference
  • Contributing
  • License

🎯 Overview

RedEval evaluates LLM safety through two complementary approaches:

🔴 Attack Phase

Tests LLM vulnerability to adversarial prompts using various jailbreaking techniques:

  • Direct attacks: Straightforward harmful prompts
  • Human jailbreaks: Human-crafted bypass techniques
  • Zero-shot attacks: Automated adversarial prompt generation

🛡️ Refuse Phase

Evaluates LLM's ability to appropriately refuse harmful requests across multiple safety datasets:

  • CoCoNot: Context-aware content moderation
  • SGXSTest: Safety guidelines examination
  • XSTest: Cross-domain safety testing
  • ORBench: Objective refusal benchmarking

📊 Scoring

Calculates comprehensive safety metrics combining both attack and refusal performance.

🏗️ Architecture

Core Components

redeval/
├── redeval/                 # Core Python modules
│   ├── config.py           # Environment & configuration management
│   ├── pipeline.py         # Centralized pipeline orchestrator
│   ├── cli.py              # Unified command-line interface
│   ├── exceptions.py       # Custom exception handling
│   ├── generate_attack.py  # Attack prompt generation
│   ├── run_attack.py       # Attack execution
│   ├── eval_attack.py      # Attack evaluation
│   ├── run_refuse.py       # Refusal testing
│   ├── eval_refuse.py      # Refusal evaluation
│   └── score.py            # Metric calculation
├── sh/                     # Shell script interfaces
│   ├── setup_env.sh        # Environment setup
│   ├── generate_attack.sh  # Attack generation script
│   ├── run_attack.sh       # Attack execution script
│   ├── eval_attack.sh      # Attack evaluation script
│   ├── run_refuse.sh       # Refusal testing script
│   ├── eval_refuse.sh      # Refusal evaluation script
│   └── score.sh            # Scoring script
├── recipes/                # Configuration files
├── logs/                   # Evaluation results
└── .env                    # Environment variables

Evaluation Pipeline

graph TD
    A[Generate Attack Prompts] --> B[Run Attack Tests]
    B --> C[Evaluate Attack Results]
    D[Run Refuse Tests] --> E[Evaluate Refuse Results]
    C --> F[Calculate Final Scores]
    E --> F
    F --> G[Generate Reports]

🛠️ Installation

Prerequisites

  • Python 3.8 or higher
  • OpenAI API key
  • HuggingFace token
  • Git

Step-by-Step Installation

  1. Clone Repository

    git clone <repository-url>
    cd redeval
    
  2. Install Dependencies

    pip install -r requirements.txt
    
  3. Configure Environment

    cp .env.template .env
    # Edit .env with your API credentials
    
  4. Verify Installation

    python -m redeval.cli --help
    

⚙️ Configuration

Environment Variables

RedEval uses environment variables for secure and flexible configuration:

Required Variables

VariableDescriptionExample
OPENAI_API_KEYOpenAI API keysk-proj-...
HUGGINGFACE_TOKENHuggingFace tokenhf_...

Optional Configuration

VariableDescriptionDefault
REDEVAL_PROJECT_ROOTProject root directoryCurrent directory
REDEVAL_LOG_DIRLogs directory./logs
REDEVAL_RECIPES_DIRConfiguration directory./recipes
REDEVAL_LOG_LEVELLogging levelINFO
REDEVAL_NUM_SAMPLESNumber of evaluation samples10
REDEVAL_SEEDRandom seed0

Model Configuration

VariableDescriptionDefault
REDEVAL_OPEN_SOURCE_MODELSOpen-source models list"Qwen/Qwen2.5-7B-Instruct"
REDEVAL_CLOSED_SOURCE_MODELSClosed-source models list"gpt-4o-mini"

Configuration Files

Configuration files in recipes/ directory control evaluation parameters:

  • attack/base-open.yml - Open-source model attack configuration
  • attack/base-close.yml - Closed-source model attack configuration
  • attack/eval.yml - Attack evaluation configuration
  • refuse/base-open.yml - Open-source model refusal configuration
  • refuse/base-close.yml - Closed-source model refusal configuration
  • refuse/eval.yml - Refusal evaluation configuration

🚀 Usage

Command Line Interface

RedEval provides a unified CLI for all operations:

Complete Pipeline

# Run full evaluation pipeline
python -m redeval.cli run-pipeline

# Run with specific models
python -m redeval.cli run-pipeline --models "Qwen/Qwen2.5-7B-Instruct" "gpt-4o-mini"

# Run specific phases only
python -m redeval.cli run-pipeline --phases generate_attack run_attack eval_attack

Individual Components

# Generate attack prompts
python -m redeval.cli generate-attack --config ./recipes/attack/base-close.yml

# Run attack evaluation
python -m redeval.cli run-attack --config ./recipes/attack/base-open.yml --model "Qwen/Qwen2.5-7B-Instruct"

# Evaluate attack results
python -m redeval.cli eval-attack --config ./recipes/attack/eval.yml --log-dir ./logs/attack/HarmBench/direct/model_name

# Run refusal tests
python -m redeval.cli run-refuse --config ./recipes/refuse/base-open.yml --model "Qwen/Qwen2.5-7B-Instruct"

# Evaluate refusal results
python -m redeval.cli eval-refuse --config ./recipes/refuse/eval.yml --log-dir ./logs/refuse/CoCoNot/base/model_name

# Calculate scores
python -m redeval.cli score --log-dir ./logs/attack/HarmBench/direct/model_name --keyword "unsafe"

Shell Script Interface

For users preferring shell scripts:

Download Tool