Skip to content
KitploitKITPLOIT
उपकरणएक्सप्लॉइटब्लॉग
Log in
जमा करें
उपकरणएक्सप्लॉइटब्लॉग
जमा करें

हैकिंग, पेनटेस्ट और साइबर सुरक्षा उपकरण आपके सुरक्षा शस्त्रागार के लिए!

Kitploit हैकिंग, साइबर सुरक्षा और पेंटेस्टिंग टूल्स की एक निर्देशिका है। कमजोरियों को खोजने, सिस्टम का विश्लेषण करने, परीक्षण को स्वचालित करने और अपनी सुरक्षा को मजबूत करने के लिए नवीनतम प्रोजेक्ट अपडेट खोजें।

फ़ीडसंपर्कगोपनीयता© 2026 Kitploit

टूल निर्देशिका

श्रेणियाँ

सभी श्रेणियाँ देखें
Loading categories
Learning-to-Detect — Detects unknown jailbreak attacks in large vision-language models using hidden state analysis and autoencoders, with training and evaluation pipelines for robust safety monitoring. | Kitploit
उपकरण/GitHubGitHub/shuangliangx/learning-to-detect
Defensive ToolsVulnerability AnalysisMachine LearningAI SecurityAnomaly Detection
GitHubshuangliangx/learning-to-detect

Learning-to-Detect

Detects unknown jailbreak attacks in large vision-language models using hidden state analysis and autoencoders, with training and evaluation pipelines for robust safety monitoring.

सबसे लोकप्रिय

सभी देखें →

हमारे समुदाय द्वारा सबसे अधिक उपयोग किए जाने वाले उपकरण खोजें।

सभी उपकरण खोजें

हमारे उपकरणों का संग्रह ब्राउज़ करें

सभी उपकरण देखें →
साझा करें
रिपॉजिटरी देखें
16821 दिन पहलेअभी तक समीक्षित नहीं
अनुरोधित भाषा में सामग्री उपलब्ध नहीं है। अंग्रेज़ी संस्करण दिखाया जा रहा है।

Learning to Detect Unknown Jailbreak Attacks in Large Vision-Language Models

Official implementation of “Learning to Detect Unknown Jailbreak Attacks in Large Vision-Language Models.”

This repository contains the data-processing, hidden-state extraction, classifier training, safety-pattern auto-encoder, and evaluation code used in the project.

This release provides the LLaVA-v1.6-Vicuna-7B implementation. Adapting the pipeline to Qwen2.5-VL or CogVLM requires model-specific input processing and separately trained detectors.

Contents

  • Models
  • Repository Structure
  • Detection Pipeline
  • Datasets

Models

The experiments use the following base models:

  • LLaVA-v1.6-Vicuna-7B, a large vision-language model based on Vicuna.
  • Llama Guard 3 8B, a safety guardrail model used to assess generated responses.

Download the model weights separately and place them under:

asset/weights/

Repository Structure

.
├── asset/
│   ├── advbench/                 # AdvBench images
│   ├── harmbench/                # HarmBench DirectRequest images
│   ├── GQA/                      # GQA images
│   ├── HiddenStates/             # Extracted hidden states
│   └── weights/                  # Model weights (not included)
├── Benchmarks/                   # Evaluation benchmark metadata
├── vicuna/
│   ├── instructions/
│   │   ├── advbench.json
│   │   ├── GQA.json
│   │   └── harmbench.json        # HarmBench DirectRequest metadata
│   ├── qa.py
│   ├── qa-baseline.py
│   └── train.py
├── autoencoder.py
├── llama3_guard.py
└── test.py

Detection Pipeline

Some scripts use paths relative to their own working directory. Run the commands from the directories shown below.

1. Query the vision-language model

Query the model on the unsafe source data (AdvBench) and safe source data (GQA):

cd vicuna
python qa.py --dataset advbench
python qa.py --dataset GQA
cd ..

The HarmBench DirectRequest data can be queried with the same interface:

cd vicuna
python qa.py --dataset harmbench
cd ..

2. Assess and split model responses

Assess the generated AdvBench responses with Llama Guard 3:

python llama3_guard.py --file vicuna/instructions/advbench.json

Create the AdvBench and GQA training/test splits:

cd vicuna/instructions
python process.py
cd ../..

3. Extract hidden states

Extract hidden states for the evaluation benchmarks currently configured in vicuna/qa-baseline.py:

cd vicuna
python qa-baseline.py
cd ..

4. Train and test the MSCAV classifiers

cd vicuna
python train.py --train
python train.py --test
cd ..

5. Train the Safety Pattern Auto-Encoder (SPAE)

python autoencoder.py

6. Evaluate detection performance

python test.py

Datasets

Training and appendix data

DatasetRoleMetadataImages
AdvBenchUnsafe source datavicuna/instructions/advbench.jsonasset/advbench/
GQASafe source datavicuna/instructions/GQA.jsonasset/GQA/
HarmBench (DirectRequest)Unsafe-source experiment reported in the appendixvicuna/instructions/harmbench.jsonasset/harmbench/

The included HarmBench DirectRequest subset contains 320 image–request pairs. Image paths in harmbench.json are repository-relative and follow the same layout convention as AdvBench.

Evaluation benchmarks

Download the evaluation images from ModelScope: detecpolo/Learning-to-Detect and copy the downloaded asset/ directory into the repository root. The matching benchmark JSON files are provided in this repository under Benchmarks/.

Please refer to Section 4.1 and the appendices of our paper for dataset sources, attack construction, and experimental settings.

टूल डाउनलोड करें