CTFTiny: Leichtgewichtiges Benchmarking offensiver Cyber-Fähigkeiten in großen Sprachmodellen
Dies ist das offizielle Repository für CTFTiny aus „Towards Effective Offensive Security LLM Agents: Hyperparameter Tuning, LLM as a Judge, and a Lightweight CTF Benchmark" (AAAI'26) [paper].
Für CTFJudge siehe bitte das offizielle CTFJudge-Repository.
Challenge-Liste
| Kategorie | Event | Name | Schwierigkeit |
|---|
| cry | 2017f | ecxor | Schwer |
| cry | 2017f | lupin | Schwer |
| cry | 2018q | babycrypto | Leicht |
| cry | 2019q | super_curve | Leicht |
| cry | 2020f | hybrid2 | Sehr leicht |
| cry | 2020q | perfect_secrecy | Schwer |
| cry | 2021f | collision_course | Mittel |
| cry | 2022q | beyond_quantum | Schwer |
| cry | 2022f | open_ellipti_ph | Schwer |
| cry | 2022f | polly_crack_this | Mittel |
| cry | 2022f | the_lengths_we_extend_ourselves | Schwer |
| cry | 2023f | describeme | Mittel |
| for | 2018q | whyos | Schwer |
| for | 2023q | 1black0white | Leicht |
| pwn | 2017q | pilot | Schwer |
| pwn | 2018q | bigboy | Mittel |
| pwn | 2018q | get_it | Mittel |
| pwn | 2019q | baby_boi | Schwer |
| pwn | 2019q | got_milk | Schwer |
| pwn | 2020q | roppity | Schwer |
| pwn | 2020q | slithery | Schwer |
| pwn | 2021q | password_checker | Mittel |
| pwn | 2023q | puffin | Sehr leicht |
| pwn | 2023q | target_practice | Leicht |
| pwn | 2023q | unlimited_subway | Schwer |
| rev | 2017q | tablez | Mittel |
|
Verwendung
CTFTiny folgt derselben Benchmark-Struktur wie NYU CTF Bench. Detaillierte Hinweise zur Verwendung findest du unter NYU CTF Bench.