
Specialized reasoning LLM for source-code vulnerability detection in C/C++ and Python, with dataset construction, SFT/DPO training, and result-reproduction scripts.
git lfs install # one-time setup
git clone https://github.com/ucsb-mlsec/VulnLLM-R.git
cd VulnLLM-R
# If you cloned before installing Git LFS, run: git lfs pull
conda create -n vulnscan python=3.11
conda activate vulnscan
pip install -e . -e ./vulscan/train/LLaMA-Factory -e ./vulscan/model_zoo
# generate VulnLLM-R-7B's results
python -m vulscan.test.test --output_dir results/test_data --dataset_path ./datasets/test/function_level/ ./datasets/test/repo_level/ --language python c java --model UCSB-SURFI/VulnLLM-R-7B --requests_per_minute 1000 --save --use_cot --batch_size 4 --tp 2 --vllm --max_tokens 8192 --random_cwe
python -m vulscan.test.test_hf \
--output_dir results/test_hf \
--hf_dataset UCSB-SURFI/VulnLLM-R-Test-Data \
--hf_split repo_level function_level \
--language c python java \
--model UCSB-SURFI/VulnLLM-R-7B \
--save --use_cot --vllm --tp 2
# [optional] generate other models' results with our shell script
# remember to add your API keys to .env file if you want to run commercial models
# use ./run_test.sh -h for more options
./vulscan/test/run_test.sh -o results/test_data -t 2 # -o means output directory, -t means tensor parallelism
./vulscan/test/run_test.sh -o results/test_data -M o3-mini # -M means model name, which runs only one model.
./vulscan/test/run_test.sh -o results/test_data -M gpt-5.4 -e high # -e sets reasoning effort (e.g., none/low/medium/high/xhigh)
./vulscan/test/run_test.sh -o results/test_data -M claude-opus-4-6 -e high
# [optional] draw plot to compare with other models
python plots/plot_language_comparison_models.py --results-dir results/test_data
python plots/plot_model_size_scatter.py --results-dir results/test_data # Note: Labels may overlap with scatter points. Adjust text positions manually if needed.
We also provide the reduced reasoning version of the distilled datasets:
Merge existing function-level vulnerability detection datasets: PrimeVul [1], SecCodePLT [2], Juliet [3], Sven [4], and Arvo [5]. Within these datasets, PrimeVul has the most complicated functions. We create two training sets: clean (without PrimeVul) and noisy (with PrimeVul), so we can train on relatively simple datasets and test on the complex PrimeVul dataset. Note that we name the training set with PrimeVul as noisy not means the dataset is noisy. It is a relatively arbitrary name we used at the beginning.
vulscan/data_process/data_utils has a set of scripts to process and merge the
datasets.
raw_to_us.py: Merge the raw data into our dataset and remove redundant datacheck_cwe_correct.py: Compute the accuracy for each CWE categorygenerate_arvo_raw_data.py: Generate structured raw data from arvo datasetarvo_to_us.py: Reformat arvo structured raw data to our dataset formatsplit_good_bad_for_juliet.py: Extract data from the raw Juliet 1.3 dataset and convert it into the required
format, which forms part of our c clean_datasetadd_sven_to_clean_dataset.py: Extract data from the Sven dataset, forming part of our C clean datasetsync_large_small.py: Synchronize the modifications of noisy_dataset/large_train/c to
noisy_dataset/small_train/cremove_testing_from_training.py: Add the human tag to each data, meaning the point has been verified by
human and used as testing data
-data_utils.py: Add the related_cwe field to dataset.datasets/clean_dataset: the training data without PrimeVul
datasets/clean_dataset/python has the data from SVEN and SecCodePLTdatasets/clean_dataset/c has the data from Juliet and SVENdatasets/noisy_dataset
datasets/noisy_dataset/small_train: Contains the training data from PrimeVul and SVEN with selected CWEs (
we use the PrimeVul data in this dataset as the training)datasets/noisy_dataset/large_train: Contains the training data from PrimeVul and SVEN and SecCodePLT with
more CWEs (This dataset can later be used to train larger models)datasets/noisy_dataset/test: A small testing set from PrimeVul verified by humandatasets/test
datasets/test/test_clean: The testing data from SVEN and SecCodePLT and Juliet; with OOD CWEs that are not
part of the training setdatasets/test/test_primevul_pair: The original PrimeVul testing datavulscan/data_process/data_utils/get_cwe_stat.py to get the histogram of the dataset| Dataset | Language | Train/test | CWE | # Benign | # Vuln. | average length |
|---|---|---|---|---|---|---|
| Clean (seccodeplt) | Python | Train | 20 | 1281 | 1281 | 741 |
| Clean (juliet) | C/C++ | Train | 22 | 1716 | 1653 | 3689 |
| Hard (primevul filtered) | C/C++ | Train | 26 | 2717 | 2952 | 4689 |
| Long Context (Oss-fuzz) | C/C++ | Train | 3 | 475 | 604 | 12761 |
| Simple (seccodeplt) | Python | Test | 24 (6 ood) | 74 | 74 | 814 |
| Simple (juliet) | C/C++ | Test | 38 (14 ood) | 358 | 376 | 2575 |
| Hard (PrimeVul, SecLLMHolmes) | C/C++ | Test | 13 (5 ood) | 145 | 152 | 4545 |
| Long Context (Oss-fuzz) | C/C++ | Test | 3 (0 ood) | 0 | 320 | 18929 |
| primevul test (noisy) | C/C++ | Test | 56 (34 ood) | 421 | 422 | 5341 |
After constructing the datasets, we will generate reasoning data for our training set.
We will query the DeepSeek-r1 and QwQ reasoning model to generate the reasoning data and filter out the ones with very
long reasoning chains.
The code for generating reasoning data is in vulscan/data_process/generate_reasoning and the reasoning data will be
saved
in datasets/reasoning_data.
cd vulscan/data_process/generate_reasoning