
vulnrepro-benchmark
Benchmark measuring AI models' ability to detect vulnerabilities in source code via real bug bounty cases with balanced recall and false-positive…
ai-securitycode-analysisctf+7
9

Benchmark measuring AI models' ability to detect vulnerabilities in source code via real bug bounty cases with balanced recall and false-positive…

Autonomous white-hat security auditor for AI-driven code review, bug bounty research, exploit construction, and execution-grounded verification.