
terminal-bench-2
Benchmark for evaluating AI agents on real-world tasks including vulnerability resolution, code debugging, and protein assembly in containerized…
ctfdevsecopseducation+5
394

Benchmark for evaluating AI agents on real-world tasks including vulnerability resolution, code debugging, and protein assembly in containerized…