Volver a actualizaciones
ActualizadaAug 31, 2026

Awesome-LLMs-for-Vulnerability-Detection — Actualizado!

El índice más completo y actualizado continuamente de la comunidad sobre investigación de Modelos de Lenguaje de Gran Tamaño para la detección de vulnerabilidades de software: artículos sobre detección a nivel de función, a nivel de repositorio, agéntica y de contratos inteligentes, además de conjuntos de datos, puntos de referencia y estudios.

Compartir

Awesome Large Language Models para Detección de Vulnerabilidades

Una lista curada de artículos, proyectos y habilidades de agentes sobre el uso de LLMs para la detección y descubrimiento de vulnerabilidades.


📄 Artículos

Solo se muestran trabajos de 2025 en adelante. Para trabajos anteriores, consulte Archivo de Artículos (2024 y anteriores).

TítuloLugarAñoArtículoGithub
VulnGym: Benchmarking Coding Agents for Repository-Level Vulnerability Detection2026enlaceenlace
VulTriage: Triple-Path Context Augmentation for LLM-Based Vulnerability Detection2026enlaceenlace
Synthesizing Multi-Agent Harnesses for Vulnerability Discovery2026enlaceenlace
QRS: A Rule-Synthesizing Neuro-Symbolic Triad for Autonomous Vulnerability Discovery2026enlace
Seclens: Role-specific Evaluation of LLM's for security vulnerablity detection2026enlaceenlace
Do Fine-Tuned LLMs Understand Vulnerabilities? An Investigation into the Semantic Trap2026enlace
Sifting the Noise: A Comparative Study of LLM Agents in Vulnerability False Positive FilteringISSTA2026enlace
AgenticSCR: An Autonomous Agentic Secure Code Review for Immature Vulnerabilities Detection2026enlace
LLM-based Vulnerability Detection at Project Scale: An Empirical Study2026enlace
MulVul: Retrieval-augmented Multi-Agent Code Vulnerability Detection via Cross-Model Prompt Evolution2026enlace
VulnLLM-R: Specialized Reasoning LLM with Agent Scaffold for Vulnerability Detection2025enlace
VULPO: Context-Aware Vulnerability Detection via On-Policy LLM Optimization2025enlace
VulInstruct: Teaching LLMs Root-Cause Reasoning for Vulnerability Detection via Security Specifications2025enlace
From Large to Mammoth: A Comparative Evaluation of Large Language Models in Vulnerability DetectionNDSS2025enlace
Benchmarking LLMs and LLM-based Agents in Practical Vulnerability Detection for Code RepositoriesACL2025enlaceenlace
A Systematic Literature Review on Detecting Software Vulnerabilities with Large Language Models2025enlaceenlace
LLMxCPG: Context-Aware Vulnerability Detection Through Code Property Graph-Guided Large Language ModelsUsenix2025enlaceenlace
CLeVeR: Multi-modal Contrastive Learning for Vulnerability Code RepresentationACL Findings2025enlaceenlace
Mono: Is Your "Clean" Vulnerability Dataset Really Solvable? Exposing and Trapping Undecidable Patches and Beyond2025enlaceenlace
Learning to Focus: Context Extraction for Efficient Code Vulnerability Detection with Language Models2025enlace
SV-TrustEval-C: Evaluating Structure and Semantic Reasoning in Large Language Models for Source Code Vulnerability AnalysisSP2025enlaceenlace
SecVulEval: Benchmarking LLMs for Real-World C/C++ Vulnerability Detection2025enlaceenlace
CVE-Bench: Benchmarking LLM-based Software Engineering Agent's Ability to Repair Real-World CVE VulnerabilitiesNAACL2025enlaceenlace
R2Vul: Learning to Reason about Software Vulnerabilities with Reinforcement Learning and Structured Reasoning Distillation2025enlaceenlace
Neuro-symbolic Static Analysis with LLM-generated Vulnerability Patterns2025enlace
Context-Enhanced Vulnerability Detection Based on Large Language Model2025enlace
Everything You Wanted to Know About LLM-based Vulnerability Detection But Were Afraid to Ask2025enlaceenlace
MOS: Towards Effective Smart Contract Vulnerability Detection through Mixture-of-Experts Tuning of Large Language Models2025enlace
Abundant Modalities Offer More Nutrients: Multi-Modal-Based Function-Level Vulnerability DetectionTOSEM2025enlaceenlace
Generative Large Language Model usage in Smart Contract Vulnerability Detection2025enlace
Closing the Gap: A User Study on the Real-world Usefulness of AI-powered Vulnerability Detection & Repair in the IDEICSE2025enlaceenlace
Vulnerability Detection with Code Language Models: How Far Are We?ICSE2025enlaceenlace
Combining Fine-Tuning and LLM-based Agents for Intuitive Smart Contract Auditing with JustificationsICSE2025enlace
LAMD: Context-driven Android Malware Detection and Classification with LLMs2025enlace
LLMs in Software Security: A Survey of Vulnerability Detection Techniques and Insights2025enlaceenlace
One-for-All Does Not Work! Enhancing Vulnerability Detection by Mixture-of-Experts (MoE)2025enlace

🚀 Proyectos

NombreDescripciónGithub
OpenAnt (Knostic)Descubrimiento de vulnerabilidades en múltiples etapas impulsado por LLM con verificación adversarialenlace
DeepAuditPlataforma de red team de IA multiagente con validación de exploits mediante sandbox Dockerenlace
AutoCVEDetección y reporte automatizado de vulnerabilidades con arquitectura multiagenteenlace
DarkmoonPlataforma autónoma de pentest con IA de código abierto (GPL-3.0) y host MCP; subagentes ofensivos por tecnología, cobertura de Active Directory y Kubernetes, más de 80 herramientas orquestadas, rastro de evidencia por hallazgoenlace
strixHerramienta autónoma de pentest con IA de código abierto instalable vía pipenlace
deepsec (Vercel)Arnés de seguridad para escaneo profundo de vulnerabilidades en codebases con agentes de codificaciónenlace

🧩 Habilidades de Agentes

NombreDescripciónEnlace
codex-security (OpenAI)Escaneo autónomo de vulnerabilidades a nivel de repositorio mediante agentes Codexenlace
defending-code-reference-harness (Anthropic)Habilidades de referencia para modelado de amenazas, escaneo, triaje y parcheo con Claude Codeenlace
security-audit-skill (Cloudflare)Habilidad de auditoría de seguridad en seis fases con agentes de búsqueda en paralelo y validación adversarialenlace

📡 arxiv.md

Captura y actualización diaria automatizada de artículos de Arxiv para palabras clave específicas mediante flujos de trabajo.

Agradecimientos

El flujo de trabajo Updated Arxiv Papers Daily del proyecto se basa en este proyecto LLM4SE. Refactoricé su código original utilizando la biblioteca arxiv.

Categorías