
terminal-bench-2
Benchmark for evaluating AI agents on real-world tasks including vulnerability resolution, code debugging, and protein assembly in containerized…
ctfdevsecopseducation+5
416

Benchmark for evaluating AI agents on real-world tasks including vulnerability resolution, code debugging, and protein assembly in containerized…

Full VAPT writeup of OWASP CICD-Goat — 9 CTFd flags captured, 4 critical + 5 high findings (incl. CVE-2024-23897) mapped to the OWASP Top 10 CI/CD…