Skip to content
KitploitKITPLOIT
도구블로그
Log in
제출
도구블로그
제출

해킹, 침투 테스트 및 사이버 보안 도구를 당신의 보안 무기고에!

Kitploit은 해킹, 사이버 보안 및 침투 테스트 도구 디렉토리입니다. 최신 프로젝트 업데이트를 발견하여 취약점을 찾고, 시스템을 분석하고, 테스트를 자동화하고, 보안을 강화하세요.

피드문의개인정보© 2026 Kitploit

도구 디렉토리

카테고리

모든 카테고리 보기
Loading categories
BoxPwnr — HackTheBox, TryHackMe, PortSwigger Labs, Cybench, picoCTF 등에서 제공하는 보안 챌린지를 대상으로 LLM과 에이전트 전략을 벤치마킹하기 위한 모듈형 프레임워크 | Kitploit
도구/GitHubGitHub/0ca/boxpwnr
Privilege EscalationReconnaissanceExploit FrameworksPayload GenerationVulnerability AnalysisWeb SecurityCTFPenetration TestingLearning & EducationAI SecurityLabs & Practice
44355242개월 전Kitploit 검토 완료

인기

모두 보기 →

커뮤니티에서 가장 많이 사용되는 도구를 찾아보세요.

모든 도구 탐색

도구 컬렉션을 둘러보세요

모든 도구 보기 →
공유
GitHub
0ca/boxpwnr

BoxPwnr

HackTheBox, TryHackMe, PortSwigger Labs, Cybench, picoCTF 등에서 제공하는 보안 챌린지를 대상으로 LLM과 에이전트 전략을 벤치마킹하기 위한 모듈형 프레임워크

저장소 보기

BoxPwnr

LLM(대규모 언어 모델)이 CTF 챌린지와 보안 랩을 스스로 얼마나 해결할 수 있는지 확인해 보는 재미있는 실험입니다. HackTheBox에서 시작해 이제 많은 플랫폼과 에이전트형 솔버를 지원합니다.

BoxPwnr은 다양한 에이전트형 아키텍처의 성능을 테스트하는 데 사용할 수 있는 플러그 앤 플레이 시스템을 제공합니다: --solver [claude_code, codex, cursor-cli, grok, kiro_cli, external, single_loop_xmltag, single_loop, single_loop_compactation, hacksynth].

지원 플랫폼: --platform [htb, htb_ctf, htb_challenges, portswigger, ctfd, local, xbow, hackbench, cybergym, exploitbench, picoctf, tryhackme, levelupctf, argus]

각 지원 플랫폼에 대한 자세한 문서는 플랫폼 구현을 참조하세요.

트레이스 및 벤치마크

모든 해결 트레이스는 BoxPwnr Traces & Benchmarks에서 확인할 수 있습니다. 각 트레이스에는 LLM의 추론, 실행된 명령어, 수신된 출력을 보여주는 전체 대화 로그가 포함되어 있습니다. 대화형 웹 뷰어에서 모든 트레이스를 재생하여 머신이 단계별로 정확히 어떻게 해결되었는지 확인할 수 있습니다.

🔬 BoxPwnr Traces & Benchmarks

Total Challenges Challenges Solved Total Traces Platforms

플랫폼해결완료율트레이스
HTB Starting Point25/25100.0%770
HTB Labs268/52651.0%783
HTB Challenges324/81839.6%732
PortSwigger Labs163/27060.4%377
XBOW102/10498.1%525
Cybench40/40100.0%2165
CyberGym476/150731.6%977
picoCTF502/50399.8%1215
TryHackMe213/47744.8%905
HackBench11/1668.8%27
ExploitBench2/424.8%58
LevelUpCTF50/25419.7%146
Argus47/6078.3%1026
BSidesSF CTF 202643/5184.3%76
Cloud Village CTF 202612/2060.0%30
Neurogrid CTF: The ultimate AI security showdown17/3647.2%197

작동 방식

BoxPwnr은 LLM(또는 Claude Code, Codex, Grok, Cursor와 같은 CLI 에이전트)을 사용하여 반복적인 프로세스를 통해 CTF / 랩 대상을 자율적으로 해결합니다:

  1. 환경: 기본적으로 명령은 Kali Linux가 포함된 Docker 컨테이너에서 실행됩니다 (--executor docker)
  • 컨테이너는 최초 실행 시 자동으로 빌드됩니다 (약 10분 소요)
  • 플랫폼에서 요구하는 경우 VPN 연결이 자동으로 설정됩니다
  1. 실행 루프 (기본 single_loop_* 솔버):
  • LLM은 작업과 제약 조건을 정의하는 상세한 시스템 프롬프트를 받습니다
  • LLM은 이전 출력을 기반으로 다음 명령을 제안합니다
  • 명령은 선택된 실행기에서 실행됩니다
  • 출력은 분석을 위해 LLM에 다시 전달됩니다
  • 플래그(또는 플랫폼 성공 기준)가 충족될 때까지 프로세스가 반복됩니다
  • CLI 기반 솔버(claude_code, codex, grok, cursor-cli, kiro_cli)는 자체 에이전트 루프를 실행하고 결과를 BoxPwnr로 스트리밍합니다
  1. 명령 자동화:
  • 에이전트는 수동 상호작용 없이 완전히 자동화된 명령을 제공하도록 지시됩니다
  • 명령에는 적절한 타임아웃이 포함되어야 하며 서비스 지연을 처리해야 합니다
  1. 결과:
  • 대화와 명령은 분석/재생을 위해 트레이스로 저장됩니다
  • 플래그를 찾으면 요약을 생성할 수 있습니다
  • 사용 통계(토큰, 비용, 턴 수)가 추적됩니다

사용 방법

사전 요구 사항

  1. 서브모듈과 함께 저장소를 클론합니다 ```bash git clone --recurse-submodules https://github.com/0ca/BoxPwnr cd BoxPwnr

    Install uv if you haven't already

    curl -LsSf https://astral.sh/uv/install.sh | sh

    Sync dependencies (creates .venv)

    uv sync

2. Docker
- BoxPwnr를 실행하려면 Docker가 설치되어 실행 중이어야 합니다
- 설치 지침은 다음에서 확인할 수 있습니다: [https://docs.docker.com/get-docker/](https://docs.docker.com/get-docker/)

### BoxPwnr 실행```bash
uv run boxpwnr --platform htb --target meow [options]

On first run, you'll be prompted for any required API keys. Keys are saved to .env for future use. CLI solvers (Claude Code, Codex, Grok, Cursor, Kiro) use their own subscription auth instead of (or in addition to) API keys.

Command Line Options

Core Options

  • --platform: Platform to use (htb, htb_ctf, htb_challenges, portswigger, ctfd, local, xbow, hackbench, cybench, cybergym, exploitbench, picoctf, tryhackme, levelupctf, argus)
  • --target: Target name (e.g., meow for HTB machine, "SQL injection UNION attack" for PortSwigger lab, or XBEN-060-24 for XBOW benchmark)
  • --debug: Enable verbose logging (shows tool names and descriptions)
  • --debug-langchain: Enable LangChain debug mode (shows full HTTP requests with tool schemas, LangChain traces, and raw API payloads - very verbose)
  • --max-turns: Maximum number of turns before stopping (e.g., --max-turns 10)
  • --max-cost: Maximum cost in USD before stopping (e.g., --max-cost 2.0)
  • --max-time: Maximum time in minutes per attempt (e.g., --max-time 60)
  • --attempts: Number of attempts to solve the target (e.g., --attempts 5 for pass@5 benchmarks)
  • --default-execution-timeout: Default timeout for command execution in seconds (default: 30)
  • --max-execution-timeout: Maximum timeout for command execution in seconds (default: 300)
  • --custom-instructions: Additional custom instructions to append to the system prompt

Platforms

  • --keep-target: Keep target (machine/lab) running after completion (useful for manual follow-up)

Analysis and Reporting

  • --analyze-attempt: Analyze failed attempts using TraceAnalyzer after completion
  • --generate-summary: Generate a solution summary after completion
  • --generate-progress: Generate a progress handoff file (progress.md) for failed/interrupted attempts. This file can be used to resume the attempt later.
  • --resume-from: Path to a progress.md file from a previous attempt. The content will be injected into the system prompt to continue from where the previous attempt left off.
  • --generate-report: Generate a new report from an existing trace directory

LLM Solver and Model Selection

도구 다운로드