
HackTheBox, TryHackMe, PortSwigger Labs, Cybench, picoCTF 등에서 제공하는 보안 챌린지를 대상으로 LLM과 에이전트 전략을 벤치마킹하기 위한 모듈형 프레임워크
LLM(대규모 언어 모델)이 CTF 챌린지와 보안 랩을 스스로 얼마나 해결할 수 있는지 확인해 보는 재미있는 실험입니다. HackTheBox에서 시작해 이제 많은 플랫폼과 에이전트형 솔버를 지원합니다.
BoxPwnr은 다양한 에이전트형 아키텍처의 성능을 테스트하는 데 사용할 수 있는 플러그 앤 플레이 시스템을 제공합니다: --solver [claude_code, codex, cursor-cli, grok, kiro_cli, external, single_loop_xmltag, single_loop, single_loop_compactation, hacksynth].
지원 플랫폼: --platform [htb, htb_ctf, htb_challenges, portswigger, ctfd, local, xbow, hackbench, cybergym, exploitbench, picoctf, tryhackme, levelupctf, argus]
각 지원 플랫폼에 대한 자세한 문서는 플랫폼 구현을 참조하세요.
모든 해결 트레이스는 BoxPwnr Traces & Benchmarks에서 확인할 수 있습니다. 각 트레이스에는 LLM의 추론, 실행된 명령어, 수신된 출력을 보여주는 전체 대화 로그가 포함되어 있습니다. 대화형 웹 뷰어에서 모든 트레이스를 재생하여 머신이 단계별로 정확히 어떻게 해결되었는지 확인할 수 있습니다.
| 플랫폼 | 해결 | 완료율 | 트레이스 |
|---|---|---|---|
| HTB Starting Point | 25/25 | 770 | |
| HTB Labs | 268/526 | 783 | |
| HTB Challenges | 324/818 | 732 | |
| PortSwigger Labs | 163/270 | 377 | |
| XBOW | 102/104 | 525 | |
| Cybench | 40/40 | 2165 | |
| CyberGym | 476/1507 | 977 | |
| picoCTF | 502/503 | 1215 | |
| TryHackMe | 213/477 | 905 | |
| HackBench | 11/16 | 27 | |
| ExploitBench | 2/42 | 58 | |
| LevelUpCTF | 50/254 | 146 | |
| Argus | 47/60 | 1026 | |
| BSidesSF CTF 2026 | 43/51 | 76 | |
| Cloud Village CTF 2026 | 12/20 | 30 | |
| Neurogrid CTF: The ultimate AI security showdown | 17/36 | 197 |
BoxPwnr은 LLM(또는 Claude Code, Codex, Grok, Cursor와 같은 CLI 에이전트)을 사용하여 반복적인 프로세스를 통해 CTF / 랩 대상을 자율적으로 해결합니다:
--executor docker)single_loop_* 솔버):claude_code, codex, grok, cursor-cli, kiro_cli)는 자체 에이전트 루프를 실행하고 결과를 BoxPwnr로 스트리밍합니다서브모듈과 함께 저장소를 클론합니다 ```bash git clone --recurse-submodules https://github.com/0ca/BoxPwnr cd BoxPwnr
curl -LsSf https://astral.sh/uv/install.sh | sh
uv sync
2. Docker
- BoxPwnr를 실행하려면 Docker가 설치되어 실행 중이어야 합니다
- 설치 지침은 다음에서 확인할 수 있습니다: [https://docs.docker.com/get-docker/](https://docs.docker.com/get-docker/)
### BoxPwnr 실행```bash
uv run boxpwnr --platform htb --target meow [options]
On first run, you'll be prompted for any required API keys. Keys are saved to .env for future use. CLI solvers (Claude Code, Codex, Grok, Cursor, Kiro) use their own subscription auth instead of (or in addition to) API keys.
--platform: Platform to use (htb, htb_ctf, htb_challenges, portswigger, ctfd, local, xbow, hackbench, cybench, cybergym, exploitbench, picoctf, tryhackme, levelupctf, argus)--target: Target name (e.g., meow for HTB machine, "SQL injection UNION attack" for PortSwigger lab, or XBEN-060-24 for XBOW benchmark)--debug: Enable verbose logging (shows tool names and descriptions)--debug-langchain: Enable LangChain debug mode (shows full HTTP requests with tool schemas, LangChain traces, and raw API payloads - very verbose)--max-turns: Maximum number of turns before stopping (e.g., --max-turns 10)--max-cost: Maximum cost in USD before stopping (e.g., --max-cost 2.0)--max-time: Maximum time in minutes per attempt (e.g., --max-time 60)--attempts: Number of attempts to solve the target (e.g., --attempts 5 for pass@5 benchmarks)--default-execution-timeout: Default timeout for command execution in seconds (default: 30)--max-execution-timeout: Maximum timeout for command execution in seconds (default: 300)--custom-instructions: Additional custom instructions to append to the system prompt--keep-target: Keep target (machine/lab) running after completion (useful for manual follow-up)--analyze-attempt: Analyze failed attempts using TraceAnalyzer after completion--generate-summary: Generate a solution summary after completion--generate-progress: Generate a progress handoff file (progress.md) for failed/interrupted attempts. This file can be used to resume the attempt later.--resume-from: Path to a progress.md file from a previous attempt. The content will be injected into the system prompt to continue from where the previous attempt left off.--generate-report: Generate a new report from an existing trace directory