
Awesome-LLMs-for-Vulnerability-Detection — Mis à jour !
L'index le plus complet et continuellement mis à jour de la communauté sur la recherche consacrée aux grands modèles de langage pour la détection de vulnérabilités logicielles — articles couvrant la détection au niveau des fonctions, au niveau des dépôts, agentique et des contrats intelligents, ainsi que des ensembles de données, des bancs d'essai et des études.
Awesome Large Language Models pour la Détection de Vulnérabilités
Une liste organisée d'articles, de projets et de compétences d'agents sur l'utilisation des LLM pour la détection et la découverte de vulnérabilités.
📄 Articles
Ne montre que 2025 et après. Pour les travaux antérieurs, voir Archives des articles (2024 et avant).
| Titre | Conférence | Année | Article | Github |
|---|---|---|---|---|
| VulnGym: Benchmarking Coding Agents for Repository-Level Vulnerability Detection | 2026 | lien | lien | |
| VulTriage: Triple-Path Context Augmentation for LLM-Based Vulnerability Detection | 2026 | lien | lien | |
| Synthesizing Multi-Agent Harnesses for Vulnerability Discovery | 2026 | lien | lien | |
| QRS: A Rule-Synthesizing Neuro-Symbolic Triad for Autonomous Vulnerability Discovery | 2026 | lien | ||
| Seclens: Role-specific Evaluation of LLM's for security vulnerablity detection | 2026 | lien | lien | |
| Do Fine-Tuned LLMs Understand Vulnerabilities? An Investigation into the Semantic Trap | 2026 | lien | ||
| Sifting the Noise: A Comparative Study of LLM Agents in Vulnerability False Positive Filtering | ISSTA | 2026 | lien | |
| AgenticSCR: An Autonomous Agentic Secure Code Review for Immature Vulnerabilities Detection | 2026 | lien | ||
| LLM-based Vulnerability Detection at Project Scale: An Empirical Study | 2026 | lien | ||
| MulVul: Retrieval-augmented Multi-Agent Code Vulnerability Detection via Cross-Model Prompt Evolution | 2026 | lien | ||
| VulnLLM-R: Specialized Reasoning LLM with Agent Scaffold for Vulnerability Detection | 2025 | lien | ||
| VULPO: Context-Aware Vulnerability Detection via On-Policy LLM Optimization | 2025 | lien | ||
| VulInstruct: Teaching LLMs Root-Cause Reasoning for Vulnerability Detection via Security Specifications | 2025 | lien | ||
| From Large to Mammoth: A Comparative Evaluation of Large Language Models in Vulnerability Detection | NDSS | 2025 | lien | |
| Benchmarking LLMs and LLM-based Agents in Practical Vulnerability Detection for Code Repositories | ACL | 2025 | lien | lien |
| A Systematic Literature Review on Detecting Software Vulnerabilities with Large Language Models | 2025 | lien | lien | |
| LLMxCPG: Context-Aware Vulnerability Detection Through Code Property Graph-Guided Large Language Models | Usenix | 2025 | lien | lien |
| CLeVeR: Multi-modal Contrastive Learning for Vulnerability Code Representation | ACL Findings | 2025 | lien | lien |
| Mono: Is Your "Clean" Vulnerability Dataset Really Solvable? Exposing and Trapping Undecidable Patches and Beyond | 2025 | lien | lien | |
| Learning to Focus: Context Extraction for Efficient Code Vulnerability Detection with Language Models | 2025 | lien | ||
| SV-TrustEval-C: Evaluating Structure and Semantic Reasoning in Large Language Models for Source Code Vulnerability Analysis | SP | 2025 | lien | lien |
| SecVulEval: Benchmarking LLMs for Real-World C/C++ Vulnerability Detection | 2025 | lien | lien | |
| CVE-Bench: Benchmarking LLM-based Software Engineering Agent's Ability to Repair Real-World CVE Vulnerabilities | NAACL | 2025 | lien | lien |
| R2Vul: Learning to Reason about Software Vulnerabilities with Reinforcement Learning and Structured Reasoning Distillation | 2025 | lien | lien | |
| Neuro-symbolic Static Analysis with LLM-generated Vulnerability Patterns | 2025 | lien | ||
| Context-Enhanced Vulnerability Detection Based on Large Language Model | 2025 | lien | ||
| Everything You Wanted to Know About LLM-based Vulnerability Detection But Were Afraid to Ask | 2025 | lien | lien | |
| MOS: Towards Effective Smart Contract Vulnerability Detection through Mixture-of-Experts Tuning of Large Language Models | 2025 | lien | ||
| Abundant Modalities Offer More Nutrients: Multi-Modal-Based Function-Level Vulnerability Detection | TOSEM | 2025 | lien | lien |
| Generative Large Language Model usage in Smart Contract Vulnerability Detection | 2025 | lien | ||
| Closing the Gap: A User Study on the Real-world Usefulness of AI-powered Vulnerability Detection & Repair in the IDE | ICSE | 2025 | lien | lien |
| Vulnerability Detection with Code Language Models: How Far Are We? | ICSE | 2025 | lien | lien |
| Combining Fine-Tuning and LLM-based Agents for Intuitive Smart Contract Auditing with Justifications | ICSE | 2025 | lien | |
| LAMD: Context-driven Android Malware Detection and Classification with LLMs | 2025 | lien | ||
| LLMs in Software Security: A Survey of Vulnerability Detection Techniques and Insights | 2025 | lien | lien | |
| One-for-All Does Not Work! Enhancing Vulnerability Detection by Mixture-of-Experts (MoE) | 2025 | lien |
🚀 Projets
| Nom | Description | Github |
|---|---|---|
| OpenAnt (Knostic) | Découverte de vulnérabilités multi-étapes propulsée par LLM avec vérification contradictoire | lien |
| DeepAudit | Plateforme d'équipe rouge IA multi-agents avec validation d'exploit par sandbox Docker | lien |
| AutoCVE | Détection et rapport automatisés de vulnérabilités avec architecture multi-agents | lien |
| Darkmoon | Plateforme de pentest IA autonome open-source (GPL-3.0) et hôte MCP ; sous-agents offensifs par technologie, couverture Active Directory et Kubernetes, plus de 80 outils orchestrés, piste de preuves par constatation | lien |
| strix | Outil de test d'intrusion IA autonome open-source installable via pip | lien |
| deepsec (Vercel) | Harnais de sécurité pour l'analyse approfondie de vulnérabilités de codebase avec agents de codage | lien |
🧩 Compétences d'agents
| Nom | Description | Lien |
|---|---|---|
| codex-security (OpenAI) | Analyse autonome de vulnérabilités au niveau du dépôt via les agents Codex | lien |
| defending-code-reference-harness (Anthropic) | Compétences de référence pour la modélisation des menaces, l'analyse, le triage et le correctif avec Claude Code | lien |
| security-audit-skill (Cloudflare) | Compétence d'audit de sécurité en six phases avec agents de chasse parallèles et validation contradictoire | lien |
📡 arxiv.md
Capture et mise à jour quotidiennes automatisées des articles Arxiv pour des mots-clés spécifiés via des workflows.
Remerciements
Le workflow Updated Arxiv Papers Daily du projet s'inspire de ce projet LLM4SE. J'ai refactorisé son code original en utilisant la bibliothèque arxiv.