
UpdatedAug 31, 2026
Awesome-LLMs-for-Vulnerability-Detection — Updated!
커뮤니티에서 가장 포괄적이고 지속적으로 업데이트되는 대규모 언어 모델(LLM) 기반 소프트웨어 취약점 탐지 연구 인덱스 — 함수 수준, 저장소 수준, 에이전트 기반, 스마트 계약 탐지에 이르는 논문과 데이터셋, 벤치마크, 설문 조사를 포함합니다.
취약점 탐지를 위한 대규모 언어 모델 모음(Awesome Large Language Models for Vulnerability Detection)
취약점 탐지 및 발견에 LLM을 활용한 논문, 프로젝트, 에이전트 스킬의 선별된 목록입니다.
📄 논문(Papers)
2025년 이후 논문만 표시합니다. 이전 연구는 논문 아카이브(2024년 및 이전)를 참조하세요.
| 제목 | 학회/저널 | 연도 | 논문 | Github |
|---|---|---|---|---|
| VulnGym: Benchmarking Coding Agents for Repository-Level Vulnerability Detection | 2026 | 링크 | 링크 | |
| VulTriage: Triple-Path Context Augmentation for LLM-Based Vulnerability Detection | 2026 | 링크 | 링크 | |
| Synthesizing Multi-Agent Harnesses for Vulnerability Discovery | 2026 | 링크 | 링크 | |
| QRS: A Rule-Synthesizing Neuro-Symbolic Triad for Autonomous Vulnerability Discovery | 2026 | 링크 | ||
| Seclens: Role-specific Evaluation of LLM's for security vulnerablity detection | 2026 | 링크 | 링크 | |
| Do Fine-Tuned LLMs Understand Vulnerabilities? An Investigation into the Semantic Trap | 2026 | 링크 | ||
| Sifting the Noise: A Comparative Study of LLM Agents in Vulnerability False Positive Filtering | ISSTA | 2026 | 링크 | |
| AgenticSCR: An Autonomous Agentic Secure Code Review for Immature Vulnerabilities Detection | 2026 | 링크 | ||
| LLM-based Vulnerability Detection at Project Scale: An Empirical Study | 2026 | 링크 | ||
| MulVul: Retrieval-augmented Multi-Agent Code Vulnerability Detection via Cross-Model Prompt Evolution | 2026 | 링크 | ||
| VulnLLM-R: Specialized Reasoning LLM with Agent Scaffold for Vulnerability Detection | 2025 | 링크 | ||
| VULPO: Context-Aware Vulnerability Detection via On-Policy LLM Optimization | 2025 | 링크 | ||
| VulInstruct: Teaching LLMs Root-Cause Reasoning for Vulnerability Detection via Security Specifications | 2025 | 링크 | ||
| From Large to Mammoth: A Comparative Evaluation of Large Language Models in Vulnerability Detection | NDSS | 2025 | 링크 | |
| Benchmarking LLMs and LLM-based Agents in Practical Vulnerability Detection for Code Repositories | ACL | 2025 | 링크 | 링크 |
| A Systematic Literature Review on Detecting Software Vulnerabilities with Large Language Models | 2025 | 링크 | 링크 | |
| LLMxCPG: Context-Aware Vulnerability Detection Through Code Property Graph-Guided Large Language Models | Usenix | 2025 | 링크 | 링크 |
| CLeVeR: Multi-modal Contrastive Learning for Vulnerability Code Representation | ACL Findings | 2025 | 링크 | 링크 |
| Mono: Is Your "Clean" Vulnerability Dataset Really Solvable? Exposing and Trapping Undecidable Patches and Beyond | 2025 | 링크 | 링크 | |
| Learning to Focus: Context Extraction for Efficient Code Vulnerability Detection with Language Models | 2025 | 링크 | ||
| SV-TrustEval-C: Evaluating Structure and Semantic Reasoning in Large Language Models for Source Code Vulnerability Analysis | SP | 2025 | 링크 | 링크 |
| SecVulEval: Benchmarking LLMs for Real-World C/C++ Vulnerability Detection | 2025 | 링크 | 링크 | |
| CVE-Bench: Benchmarking LLM-based Software Engineering Agent's Ability to Repair Real-World CVE Vulnerabilities | NAACL | 2025 | 링크 | 링크 |
| R2Vul: Learning to Reason about Software Vulnerabilities with Reinforcement Learning and Structured Reasoning Distillation | 2025 | 링크 | 링크 | |
| Neuro-symbolic Static Analysis with LLM-generated Vulnerability Patterns | 2025 | 링크 | ||
| Context-Enhanced Vulnerability Detection Based on Large Language Model | 2025 | 링크 | ||
| Everything You Wanted to Know About LLM-based Vulnerability Detection But Were Afraid to Ask | 2025 | 링크 | 링크 | |
| MOS: Towards Effective Smart Contract Vulnerability Detection through Mixture-of-Experts Tuning of Large Language Models | 2025 | 링크 | ||
| Abundant Modalities Offer More Nutrients: Multi-Modal-Based Function-Level Vulnerability Detection | TOSEM | 2025 | 링크 | 링크 |
| Generative Large Language Model usage in Smart Contract Vulnerability Detection | 2025 | 링크 | ||
| Closing the Gap: A User Study on the Real-world Usefulness of AI-powered Vulnerability Detection & Repair in the IDE | ICSE | 2025 | 링크 | 링크 |
| Vulnerability Detection with Code Language Models: How Far Are We? | ICSE | 2025 | 링크 | 링크 |
| Combining Fine-Tuning and LLM-based Agents for Intuitive Smart Contract Auditing with Justifications | ICSE | 2025 | 링크 | |
| LAMD: Context-driven Android Malware Detection and Classification with LLMs | 2025 | 링크 | ||
| LLMs in Software Security: A Survey of Vulnerability Detection Techniques and Insights | 2025 | 링크 | 링크 | |
| One-for-All Does Not Work! Enhancing Vulnerability Detection by Mixture-of-Experts (MoE) | 2025 | 링크 |
🚀 프로젝트(Projects)
| 이름 | 설명 | Github |
|---|---|---|
| OpenAnt (Knostic) | 적대적 검증을 통한 LLM 기반 다단계 취약점 발견 | 링크 |
| DeepAudit | Docker 샌드박스 익스플로잇 검증을 갖춘 멀티 에이전트 AI 레드팀 플랫폼 | 링크 |
| AutoCVE | 멀티 에이전트 아키텍처를 통한 자동화된 취약점 탐지 및 보고 | 링크 |
| Darkmoon | 오픈소스(GPL-3.0) 자율 AI 침투 테스트 플랫폼 및 MCP 호스트; 기술별 공격 하위 에이전트, Active Directory 및 Kubernetes 지원, 80개 이상의 오케스트레이션 도구, 발견 항목별 증거 추적 | 링크 |
| strix | pip을 통한 오픈소스 자율 AI 침투 테스트 도구 | 링크 |
| deepsec (Vercel) | 코딩 에이전트를 통한 심층 코드베이스 취약점 스캐닝용 보안 하네스 | 링크 |
🧩 에이전트 스킬(Agent Skills)
| 이름 | 설명 | 링크 |
|---|---|---|
| codex-security (OpenAI) | Codex 에이전트를 통한 자율 저장소 수준 취약점 스캐닝 | 링크 |
| defending-code-reference-harness (Anthropic) | Claude Code를 사용한 위협 모델링, 스캐닝, 트리아지, 패치를 위한 참조 스킬 | 링크 |
| security-audit-skill (Cloudflare) | 병렬 헌팅 에이전트와 적대적 검증을 갖춘 6단계 보안 감사 스킬 | 링크 |
📡 arxiv.md
워크플로우를 통해 지정된 키워드에 대한 Arxiv 논문을 매일 자동으로 수집하고 업데이트합니다.
감사의 말(Acknowledgements)
이 프로젝트의 Updated Arxiv Papers Daily 워크플로우는 LLM4SE 프로젝트에서 차용했습니다. arxiv 라이브러리를 사용하여 원본 코드를 리팩토링했습니다.