
terminal-bench-2
Benchmark for evaluating AI agents on real-world tasks including vulnerability resolution, code debugging, and protein assembly in containerized…
ctfdevsecopseducation+5

Benchmark for evaluating AI agents on real-world tasks including vulnerability resolution, code debugging, and protein assembly in containerized…