Skip to content
KitploitKITPLOIT
ToolsBlog
Log in
Submit
ToolsBlog
Submit

Hacking, PenTest, and Cybersecurity Tools for Your Security Arsenal!

Kitploit is a directory of hacking, cybersecurity, and pentesting tools. Discover the latest project updates to find vulnerabilities, analyze systems, automate testing, and strengthen your security.

··Feeds·Contact·Privacy·© 2026 Kitploit

Tool Directory

Categories

View all categories
Loading categories
VulnLLM-R — Specialized reasoning LLM for source-code vulnerability detection in C/C++ and Python, with dataset construction, SFT/DPO training, and result-reproduction scripts. | Kitploit
Tools/GitHubGitHub/ucsb-mlsec/vulnllm-r
Static Code Analysis (SAST)Vulnerability AnalysisCode AnalysisMachine LearningPapers & Research
GitHubucsb-mlsec/vulnllm-r

VulnLLM-R

Specialized reasoning LLM for source-code vulnerability detection in C/C++ and Python, with dataset construction, SFT/DPO training, and result-reproduction scripts.

View Repository
37058236 months agoReviewed by Kitploit

Most Popular

View all →

Discover the most used tools by our community.

Explore all tools

Browse our collection of tools

View all tools →
Share

VulnLLM-R: Specialized Reasoning LLM for Vulnerability Detection

  • Paper: arXiv:2512.07533
  • Code & Data: GitHub
  • Demo: Web demo
  • Model: 7B Model
model_size_vs_f1_scatter_01

Environment and dataset

🛠️ Create environment

  • Install Git LFS and clone the repository (LFS files are fetched automatically during clone)
git lfs install          # one-time setup
git clone https://github.com/ucsb-mlsec/VulnLLM-R.git
cd VulnLLM-R
# If you cloned before installing Git LFS, run: git lfs pull
  • Create a new conda environment
conda create -n vulnscan python=3.11
conda activate vulnscan
  • Install the required packages
pip install -e . -e ./vulscan/train/LLaMA-Factory -e ./vulscan/model_zoo

For Reproducing Our Results

# generate VulnLLM-R-7B's results
python -m vulscan.test.test --output_dir results/test_data --dataset_path ./datasets/test/function_level/ ./datasets/test/repo_level/ --language python c java --model UCSB-SURFI/VulnLLM-R-7B --requests_per_minute 1000 --save --use_cot --batch_size 4 --tp 2 --vllm --max_tokens 8192 --random_cwe

python -m vulscan.test.test_hf \
      --output_dir results/test_hf \
      --hf_dataset UCSB-SURFI/VulnLLM-R-Test-Data \
      --hf_split repo_level function_level \
      --language c python java \
      --model UCSB-SURFI/VulnLLM-R-7B \
      --save --use_cot --vllm --tp 2

# [optional] generate other models' results with our shell script 
# remember to add your API keys to .env file if you want to run commercial models
# use ./run_test.sh -h for more options
./vulscan/test/run_test.sh -o results/test_data -t 2 # -o means output directory, -t means tensor parallelism
./vulscan/test/run_test.sh -o results/test_data -M o3-mini # -M means model name, which runs only one model.
./vulscan/test/run_test.sh -o results/test_data -M gpt-5.4 -e high # -e sets reasoning effort (e.g., none/low/medium/high/xhigh)
./vulscan/test/run_test.sh -o results/test_data -M claude-opus-4-6 -e high

# [optional] draw plot to compare with other models
python plots/plot_language_comparison_models.py --results-dir results/test_data
python plots/plot_model_size_scatter.py --results-dir results/test_data # Note: Labels may overlap with scatter points. Adjust text positions manually if needed.

Existing Distilled Datasets

  • Distill-DeepSeek
  • Distill-QwQ

We also provide the reduced reasoning version of the distilled datasets:

  • Reduced-Distill-DeepSeek
  • Reduced-Distill-QwQ

Technical Details

📚 Construct training and testing datasets

Merge existing function-level vulnerability detection datasets: PrimeVul [1], SecCodePLT [2], Juliet [3], Sven [4], and Arvo [5]. Within these datasets, PrimeVul has the most complicated functions. We create two training sets: clean (without PrimeVul) and noisy (with PrimeVul), so we can train on relatively simple datasets and test on the complex PrimeVul dataset. Note that we name the training set with PrimeVul as noisy not means the dataset is noisy. It is a relatively arbitrary name we used at the beginning.

  • After download all the datasets, vulscan/data_process/data_utils has a set of scripts to process and merge the datasets.
    • raw_to_us.py: Merge the raw data into our dataset and remove redundant data
    • check_cwe_correct.py: Compute the accuracy for each CWE category
    • generate_arvo_raw_data.py: Generate structured raw data from arvo dataset
    • arvo_to_us.py: Reformat arvo structured raw data to our dataset format
    • split_good_bad_for_juliet.py: Extract data from the raw Juliet 1.3 dataset and convert it into the required format, which forms part of our c clean_dataset
    • add_sven_to_clean_dataset.py: Extract data from the Sven dataset, forming part of our C clean dataset
    • sync_large_small.py: Synchronize the modifications of noisy_dataset/large_train/c to noisy_dataset/small_train/c
    • remove_testing_from_training.py: Add the human tag to each data, meaning the point has been verified by human and used as testing data -data_utils.py: Add the related_cwe field to dataset.
  • The merged data will be saved in
    • datasets/clean_dataset: the training data without PrimeVul
      • datasets/clean_dataset/python has the data from SVEN and SecCodePLT
      • datasets/clean_dataset/c has the data from Juliet and SVEN
    • datasets/noisy_dataset
      • datasets/noisy_dataset/small_train: Contains the training data from PrimeVul and SVEN with selected CWEs ( we use the PrimeVul data in this dataset as the training)
      • datasets/noisy_dataset/large_train: Contains the training data from PrimeVul and SVEN and SecCodePLT with more CWEs (This dataset can later be used to train larger models)
      • datasets/noisy_dataset/test: A small testing set from PrimeVul verified by human
    • datasets/test
      • datasets/test/test_clean: The testing data from SVEN and SecCodePLT and Juliet; with OOD CWEs that are not part of the training set
      • datasets/test/test_primevul_pair: The original PrimeVul testing data
  • Dataset statistics; can run vulscan/data_process/data_utils/get_cwe_stat.py to get the histogram of the dataset
DatasetLanguageTrain/testCWE# Benign# Vuln.average length
Clean (seccodeplt)PythonTrain2012811281741
Clean (juliet)C/C++Train22171616533689
Hard (primevul filtered)C/C++Train26271729524689
Long Context (Oss-fuzz)C/C++Train347560412761
Simple (seccodeplt)PythonTest24 (6 ood)7474814
Simple (juliet)C/C++Test38 (14 ood)3583762575
Hard (PrimeVul, SecLLMHolmes)C/C++Test13 (5 ood)1451524545
Long Context (Oss-fuzz)C/C++Test3 (0 ood)032018929
primevul test (noisy)C/C++Test56 (34 ood)4214225341

🤔 Generate reasoning data for training

After constructing the datasets, we will generate reasoning data for our training set. We will query the DeepSeek-r1 and QwQ reasoning model to generate the reasoning data and filter out the ones with very long reasoning chains. The code for generating reasoning data is in vulscan/data_process/generate_reasoning and the reasoning data will be saved in datasets/reasoning_data.

cd vulscan/data_process/generate_reasoning
Download Tool