Skip to content
KitploitKITPLOIT
ToolsExploitsBlog
Log in
Einreichen
ToolsExploitsBlog
Einreichen

Hacking-, PenTest- und Cybersicherheits-Tools für Ihr Sicherheitsarsenal!

Kitploit ist ein Verzeichnis von Hacking-, Cybersicherheits- und Pentesting-Tools. Entdecken Sie die neuesten Projekt-Updates, um Schwachstellen zu finden, Systeme zu analysieren, Tests zu automatisieren und Ihre Sicherheit zu stärken.

··Feeds·Kontakt·Datenschutz·© 2026 Kitploit

Tool-Verzeichnis

Kategorien

Alle Kategorien anzeigen
Loading categories
BoxPwnr — A modular framework for benchmarking LLMs and agentic strategies on security challenges across HackTheBox, TryHackMe, PortSwigger Labs, Cybench, picoCTF and more. | Kitploit
Tools/GitHubGitHub/0ca/boxpwnr
Privilege EscalationReconnaissanceExploit FrameworksPayload GenerationVulnerability AnalysisWeb SecurityCTFPenetration TestingLearning & EducationAI SecurityLabs & Practice
4435524vor 2 MonatenVon Kitploit geprüft

Beliebteste

Alle anzeigen →

Entdecken Sie die meistgenutzten Tools unserer Community.

Alle Tools erkunden

Durchsuchen Sie unsere Tool-Sammlung

Alle Tools anzeigen →
Teilen
GitHub
0ca/boxpwnr

BoxPwnr

A modular framework for benchmarking LLMs and agentic strategies on security challenges across HackTheBox, TryHackMe, PortSwigger Labs, Cybench, picoCTF and more.

Repository anzeigen
Inhalt in der angeforderten Sprache nicht verfügbar. Englische Version wird angezeigt.

BoxPwnr

A fun experiment to see how far Large Language Models (LLMs) can go in solving CTF challenges and security labs on their own. It started with HackTheBox and now covers many platforms and agentic solvers.

BoxPwnr provides a plug and play system that can be used to test performance of different agentic architectures: --solver [claude_code, codex, cursor-cli, grok, kiro_cli, external, single_loop_xmltag, single_loop, single_loop_compactation, hacksynth].

Supported platforms: --platform [htb, htb_ctf, htb_challenges, portswigger, ctfd, local, xbow, hackbench, cybench, cybergym, exploitbench, picoctf, tryhackme, levelupctf, argus]

See Platform Implementations for detailed documentation on each supported platform.

Traces & Benchmarks

All solving traces are available in BoxPwnr Traces & Benchmarks. Each trace includes full conversation logs showing LLM reasoning, commands executed, and outputs received. You can replay any trace in an interactive web viewer to see exactly how the machine was solved step-by-step.

🔬 BoxPwnr Traces & Benchmarks

Total Challenges Challenges Solved Total Traces Platforms

PlatformSolvedCompletionTraces
HTB Starting Point25/25100.0%770
HTB Labs268/52651.0%783
HTB Challenges324/81839.6%732
PortSwigger Labs163/27060.4%377
XBOW102/10498.1%525
Cybench40/40100.0%2165
CyberGym476/150731.6%977
picoCTF502/50399.8%1215
TryHackMe213/47744.8%905
HackBench11/1668.8%27
ExploitBench2/424.8%58
LevelUpCTF50/25419.7%146
Argus47/6078.3%1026
BSidesSF CTF 202643/5184.3%76
Cloud Village CTF 202612/2060.0%30
Neurogrid CTF: The ultimate AI security showdown17/3647.2%197

How it Works

BoxPwnr uses LLMs (or CLI agents such as Claude Code, Codex, Grok, or Cursor) to autonomously solve CTF / lab targets through an iterative process:

  1. Environment: By default, commands run in a Docker container with Kali Linux (--executor docker)
  • Container is automatically built on first run (takes ~10 minutes)
  • VPN connection is automatically established when the platform requires it
  1. Execution Loop (default single_loop_* solvers):
  • LLM receives a detailed system prompt that defines its task and constraints
  • LLM suggests next command based on previous outputs
  • Command is executed in the chosen executor
  • Output is fed back to LLM for analysis
  • Process repeats until the flag (or platform success criteria) is met
  • CLI-based solvers (claude_code, codex, grok, cursor-cli, kiro_cli) run their own agent loop and stream results back to BoxPwnr
  1. Command Automation:
  • Agents are instructed to provide fully automated commands with no manual interaction
  • Commands should include proper timeouts and handle service delays
  1. Results:
  • Conversation and commands are saved as traces for analysis / replay
  • Summary can be generated when a flag is found
  • Usage statistics (tokens, cost, turns) are tracked

Usage

Prerequisites

  1. Clone the repository with submodules
 git clone --recurse-submodules https://github.com/0ca/BoxPwnr
 cd BoxPwnr

 # Install uv if you haven't already
 curl -LsSf https://astral.sh/uv/install.sh | sh

 # Sync dependencies (creates .venv)
 uv sync
  1. Docker
  • BoxPwnr requires Docker to be installed and running
  • Installation instructions can be found at: https://docs.docker.com/get-docker/

Run BoxPwnr

uv run boxpwnr --platform htb --target meow [options]

On first run, you'll be prompted for any required API keys. Keys are saved to .env for future use. CLI solvers (Claude Code, Codex, Grok, Cursor, Kiro) use their own subscription auth instead of (or in addition to) API keys.

Command Line Options

Core Options

Tool herunterladen