アップデート一覧に戻る
UpdatedAug 31, 2026

Awesome-LLMs-for-Vulnerability-Detection — Updated!

コミュニティで最も包括的かつ継続的に更新される、ソフトウェア脆弱性検出のための大規模言語モデルに関する研究インデックス — 関数レベル、リポジトリレベル、エージェント型、スマートコントラクト検出にわたる論文に加え、データセット、ベンチマーク、サーベイを収録。

共有

脆弱性検出のための大規模言語モデル(Awesome LLMs for Vulnerability Detection)

脆弱性の検出と発見にLLMを活用した論文、プロジェクト、エージェントスキルの厳選リスト。


📄 論文

2025年以降のみを表示しています。それ以前の研究については、論文アーカイブ(2024年以前)を参照してください。

タイトル会議/ジャーナル論文Github
VulnGym: Benchmarking Coding Agents for Repository-Level Vulnerability Detection2026linklink
VulTriage: Triple-Path Context Augmentation for LLM-Based Vulnerability Detection2026linklink
Synthesizing Multi-Agent Harnesses for Vulnerability Discovery2026linklink
QRS: A Rule-Synthesizing Neuro-Symbolic Triad for Autonomous Vulnerability Discovery2026link
Seclens: Role-specific Evaluation of LLM's for security vulnerablity detection2026linklink
Do Fine-Tuned LLMs Understand Vulnerabilities? An Investigation into the Semantic Trap2026link
Sifting the Noise: A Comparative Study of LLM Agents in Vulnerability False Positive FilteringISSTA2026link
AgenticSCR: An Autonomous Agentic Secure Code Review for Immature Vulnerabilities Detection2026link
LLM-based Vulnerability Detection at Project Scale: An Empirical Study2026link
MulVul: Retrieval-augmented Multi-Agent Code Vulnerability Detection via Cross-Model Prompt Evolution2026link
VulnLLM-R: Specialized Reasoning LLM with Agent Scaffold for Vulnerability Detection2025link
VULPO: Context-Aware Vulnerability Detection via On-Policy LLM Optimization2025link
VulInstruct: Teaching LLMs Root-Cause Reasoning for Vulnerability Detection via Security Specifications2025link
From Large to Mammoth: A Comparative Evaluation of Large Language Models in Vulnerability DetectionNDSS2025link
Benchmarking LLMs and LLM-based Agents in Practical Vulnerability Detection for Code RepositoriesACL2025linklink
A Systematic Literature Review on Detecting Software Vulnerabilities with Large Language Models2025linklink
LLMxCPG: Context-Aware Vulnerability Detection Through Code Property Graph-Guided Large Language ModelsUsenix2025linklink
CLeVeR: Multi-modal Contrastive Learning for Vulnerability Code RepresentationACL Findings2025linklink
Mono: Is Your "Clean" Vulnerability Dataset Really Solvable? Exposing and Trapping Undecidable Patches and Beyond2025linklink
Learning to Focus: Context Extraction for Efficient Code Vulnerability Detection with Language Models2025link
SV-TrustEval-C: Evaluating Structure and Semantic Reasoning in Large Language Models for Source Code Vulnerability AnalysisSP2025linklink
SecVulEval: Benchmarking LLMs for Real-World C/C++ Vulnerability Detection2025linklink
CVE-Bench: Benchmarking LLM-based Software Engineering Agent's Ability to Repair Real-World CVE VulnerabilitiesNAACL2025linklink
R2Vul: Learning to Reason about Software Vulnerabilities with Reinforcement Learning and Structured Reasoning Distillation2025linklink
Neuro-symbolic Static Analysis with LLM-generated Vulnerability Patterns2025link
Context-Enhanced Vulnerability Detection Based on Large Language Model2025link
Everything You Wanted to Know About LLM-based Vulnerability Detection But Were Afraid to Ask2025linklink
MOS: Towards Effective Smart Contract Vulnerability Detection through Mixture-of-Experts Tuning of Large Language Models2025link
Abundant Modalities Offer More Nutrients: Multi-Modal-Based Function-Level Vulnerability DetectionTOSEM2025linklink
Generative Large Language Model usage in Smart Contract Vulnerability Detection2025link
Closing the Gap: A User Study on the Real-world Usefulness of AI-powered Vulnerability Detection & Repair in the IDEICSE2025linklink
Vulnerability Detection with Code Language Models: How Far Are We?ICSE2025linklink
Combining Fine-Tuning and LLM-based Agents for Intuitive Smart Contract Auditing with JustificationsICSE2025link
LAMD: Context-driven Android Malware Detection and Classification with LLMs2025link
LLMs in Software Security: A Survey of Vulnerability Detection Techniques and Insights2025linklink
One-for-All Does Not Work! Enhancing Vulnerability Detection by Mixture-of-Experts (MoE)2025link

🚀 プロジェクト

名前説明Github
OpenAnt (Knostic)敵対的検証を備えたLLM駆動の多段階脆弱性発見link
DeepAuditDockerサンドボックスによるエクスプロイト検証を備えたマルチエージェントAIレッドチームプラットフォームlink
AutoCVEマルチエージェントアーキテクチャによる自動脆弱性検出とレポート作成link
Darkmoonオープンソース(GPL-3.0)の自律型AIペンテストプラットフォーム兼MCPホスト。技術別の攻撃用サブエージェント、Active DirectoryおよびKubernetes対応、80以上のオーケストレーション済みツール、発見事項ごとの証跡を備えるlink
strixpip経由で利用可能なオープンソースの自律型AIペネトレーションテストツールlink
deepsec (Vercel)コーディングエージェントによる大規模コードベースの脆弱性スキャンのためのセキュリティハーネスlink

🧩 エージェントスキル

名前説明リンク
codex-security (OpenAI)Codexエージェントによるリポジトリレベルの自律型脆弱性スキャンlink
defending-code-reference-harness (Anthropic)Claude Codeによる脅威モデリング、スキャン、トリアージ、パッチ適用のためのリファレンススキルlink
security-audit-skill (Cloudflare)並列ハンティングエージェントと敵対的検証を備えた6フェーズのセキュリティ監査スキルlink

📡 arxiv.md

ワークフローを通じて指定されたキーワードに関するArxiv論文を毎日自動的に取得・更新します。

謝辞

本プロジェクトのUpdated Arxiv Papers Dailyワークフローは、このプロジェクトLLM4SEを参考にしています。元のコードをarxivライブラリを使用してリファクタリングしました。

カテゴリ