
Awesome-LLMs-for-Vulnerability-Detection — Updated!
L'indice più completo e costantemente aggiornato della community sulla ricerca relativa ai Large Language Models per il rilevamento delle vulnerabilità software — articoli su rilevamento a livello di funzione, a livello di repository, agentico e di smart contract, oltre a dataset, benchmark e survey.
Awesome Large Language Models per il Rilevamento delle Vulnerabilità
Una lista curata di articoli, progetti e competenze degli agenti sull'uso degli LLM per il rilevamento e la scoperta delle vulnerabilità.
📄 Articoli
Vengono mostrati solo gli articoli del 2025 e successivi. Per i lavori precedenti, consulta Archivio Articoli (2024 e precedenti).
| Titolo | Sede | Anno | Articolo | Github |
|---|---|---|---|---|
| VulnGym: Benchmarking Coding Agents for Repository-Level Vulnerability Detection | 2026 | link | link | |
| VulTriage: Triple-Path Context Augmentation for LLM-Based Vulnerability Detection | 2026 | link | link | |
| Synthesizing Multi-Agent Harnesses for Vulnerability Discovery | 2026 | link | link | |
| QRS: A Rule-Synthesizing Neuro-Symbolic Triad for Autonomous Vulnerability Discovery | 2026 | link | ||
| Seclens: Role-specific Evaluation of LLM's for security vulnerablity detection | 2026 | link | link | |
| Do Fine-Tuned LLMs Understand Vulnerabilities? An Investigation into the Semantic Trap | 2026 | link | ||
| Sifting the Noise: A Comparative Study of LLM Agents in Vulnerability False Positive Filtering | ISSTA | 2026 | link | |
| AgenticSCR: An Autonomous Agentic Secure Code Review for Immature Vulnerabilities Detection | 2026 | link | ||
| LLM-based Vulnerability Detection at Project Scale: An Empirical Study | 2026 | link | ||
| MulVul: Retrieval-augmented Multi-Agent Code Vulnerability Detection via Cross-Model Prompt Evolution | 2026 | link | ||
| VulnLLM-R: Specialized Reasoning LLM with Agent Scaffold for Vulnerability Detection | 2025 | link | ||
| VULPO: Context-Aware Vulnerability Detection via On-Policy LLM Optimization | 2025 | link | ||
| VulInstruct: Teaching LLMs Root-Cause Reasoning for Vulnerability Detection via Security Specifications | 2025 | link | ||
| From Large to Mammoth: A Comparative Evaluation of Large Language Models in Vulnerability Detection | NDSS | 2025 | link | |
| Benchmarking LLMs and LLM-based Agents in Practical Vulnerability Detection for Code Repositories | ACL | 2025 | link | link |
| A Systematic Literature Review on Detecting Software Vulnerabilities with Large Language Models | 2025 | link | link | |
| LLMxCPG: Context-Aware Vulnerability Detection Through Code Property Graph-Guided Large Language Models | Usenix | 2025 | link | link |
| CLeVeR: Multi-modal Contrastive Learning for Vulnerability Code Representation | ACL Findings | 2025 | link | link |
| Mono: Is Your "Clean" Vulnerability Dataset Really Solvable? Exposing and Trapping Undecidable Patches and Beyond | 2025 | link | link | |
| Learning to Focus: Context Extraction for Efficient Code Vulnerability Detection with Language Models | 2025 | link | ||
| SV-TrustEval-C: Evaluating Structure and Semantic Reasoning in Large Language Models for Source Code Vulnerability Analysis | SP | 2025 | link | link |
| SecVulEval: Benchmarking LLMs for Real-World C/C++ Vulnerability Detection | 2025 | link | link | |
| CVE-Bench: Benchmarking LLM-based Software Engineering Agent's Ability to Repair Real-World CVE Vulnerabilities | NAACL | 2025 | link | link |
| R2Vul: Learning to Reason about Software Vulnerabilities with Reinforcement Learning and Structured Reasoning Distillation | 2025 | link | link | |
| Neuro-symbolic Static Analysis with LLM-generated Vulnerability Patterns | 2025 | link | ||
| Context-Enhanced Vulnerability Detection Based on Large Language Model | 2025 | link | ||
| Everything You Wanted to Know About LLM-based Vulnerability Detection But Were Afraid to Ask | 2025 | link | link | |
| MOS: Towards Effective Smart Contract Vulnerability Detection through Mixture-of-Experts Tuning of Large Language Models | 2025 | link | ||
| Abundant Modalities Offer More Nutrients: Multi-Modal-Based Function-Level Vulnerability Detection | TOSEM | 2025 | link | link |
| Generative Large Language Model usage in Smart Contract Vulnerability Detection | 2025 | link | ||
| Closing the Gap: A User Study on the Real-world Usefulness of AI-powered Vulnerability Detection & Repair in the IDE | ICSE | 2025 | link | link |
| Vulnerability Detection with Code Language Models: How Far Are We? | ICSE | 2025 | link | link |
| Combining Fine-Tuning and LLM-based Agents for Intuitive Smart Contract Auditing with Justifications | ICSE | 2025 | link | |
| LAMD: Context-driven Android Malware Detection and Classification with LLMs | 2025 | link | ||
| LLMs in Software Security: A Survey of Vulnerability Detection Techniques and Insights | 2025 | link | link | |
| One-for-All Does Not Work! Enhancing Vulnerability Detection by Mixture-of-Experts (MoE) | 2025 | link |
🚀 Progetti
| Nome | Descrizione | Github |
|---|---|---|
| OpenAnt (Knostic) | Scoperta di vulnerabilità multi-fase basata su LLM con verifica avversaria | link |
| DeepAudit | Piattaforma multi-agente di red team AI con validazione degli exploit tramite sandbox Docker | link |
| AutoCVE | Rilevamento e segnalazione automatizzata delle vulnerabilità con architettura multi-agente | link |
| Darkmoon | Piattaforma open-source (GPL-3.0) autonoma di pentest AI e host MCP; sotto-agenti offensivi per tecnologia, copertura di Active Directory e Kubernetes, oltre 80 strumenti orchestrati, traccia delle evidenze per ogni scoperta | link |
| strix | Strumento open-source autonomo di penetration testing AI installabile tramite pip | link |
| deepsec (Vercel) | Struttura di sicurezza per la scansione approfondita delle vulnerabilità del codebase con agenti di codifica | link |
🧩 Competenze degli Agenti
| Nome | Descrizione | Link |
|---|---|---|
| codex-security (OpenAI) | Scansione autonoma delle vulnerabilità a livello di repository tramite agenti Codex | link |
| defending-code-reference-harness (Anthropic) | Competenze di riferimento per threat modeling, scansione, triage e correzione con Claude Code | link |
| security-audit-skill (Cloudflare) | Competenza di audit di sicurezza in sei fasi con agenti di ricerca paralleli e validazione avversaria | link |
📡 arxiv.md
Acquisizione e aggiornamento automatico giornaliero degli articoli Arxiv per parole chiave specifiche tramite workflow.
Ringraziamenti
Il workflow Updated Arxiv Papers Daily del progetto si basa su questo progetto LLM4SE. Ho rifattorizzato il suo codice originale utilizzando la libreria arxiv.