Skip to content
KitploitKITPLOIT
ToolsExploitsBlog
Log in
Submit
ToolsExploitsBlog
Submit

Hacking, PenTest, and Cybersecurity Tools for Your Security Arsenal!

Kitploit is a directory of hacking, cybersecurity, and pentesting tools. Discover the latest project updates to find vulnerabilities, analyze systems, automate testing, and strengthen your security.

FeedsContactPrivacy© 2026 Kitploit

Tool Directory

Categories

View all categories
Loading categories
mulot — Agentic AI web pentester that drives a browser | Kitploit
Tools/GitHubGitHub/io-tl/mulot
ReconnaissanceVulnerability ScannersPayload GenerationWeb Proxies & InterceptionWeb Application ExploitationWeb SecurityFuzzingPenetration TestingLearning & EducationAI SecurityRepository Deleted
GitHubio-tl/mulot

mulot

Agentic AI web pentester that drives a browser

The upstream repository was not found during the latest Kitploit update check. This listing remains available for reference, but it has been removed from search results.
9143 months agoNot yet reviewed

Most Popular

View all →

Discover the most used tools by our community.

Explore all tools

Browse our collection of tools

View all tools →
Share

mulot

Go MCP Drives License

Agentic AI web pentester that drives a browser.

An open-weights LLM (GLM-5.2, Gemma or Qwen) drives a real headless Chromium through a Burp-style toolkit and works a target the way a human pentester would.

No frontier model, no agent running inside a Kali VM.

mulot solving OverTheWire natas16

recon → writes its own JavaScript payload → command injection → flag.

Running only this harness, open-weights GLM-5.2 agents solved 87% of OverTheWire Natas and 73% of Root-Me Web-Server (full tables below).

  • Real browser, not just HTTP real login flows, JS-heavy apps, and DOM-based XSS.
  • Burp-shaped primitives traffic history, repeater, intruder, passive scan.
  • No frontier API, no agent-in-a-VM Qwen 3.6 27B is workable; GLM-5.2 is the sweet spot.

The idea

An agent isn't just a model, the tooling often matters more than the parameter count.

Models have read the security literature, intercepting proxies, request history, site maps, fuzzing an insertion point. Burp, ZAP, mitmproxy, Fiddler their vocabulary and their gestures exists in the corpus the model was trained on.

So mulot shapes the harness to expose exactly those primitives.

Two halves

mulot splits in two: a proxy that captures and replays every exchange, and a thinking layer that runs the agent's own code inside the target page.

The proxy half

  • Traffic journal (History): every HTTP exchange lands in an always-on SQLite database, queryable and replayable. Capture goes through the browser's CDP protocol, so HTTPS is read already decrypted no interception certificate, no real MITM.

  • Request editor (Interceptor/Repeater): rebuild a request from a URL or reseed one from a captured flow, tamper with it, and reissue it, outside the browser's CORS rules.

  • Http fuzz (Intruder): one marked insertion point, a payload set swapped in turn, and match conditions on status, length, or regex.

  • Scans (Passive scan): passive and active passes over that same journal and the live DOM.

The thinking half

  • In-page JavaScript: a JS toolbox run inside the page, not on the host. The agent automates from within, the machine that runs it stays out of reach. It injects helper JS libraries into the DOM context and creates its own JavaScript tools for padding-oracle attacks, time-based SQLi, deserialization...

  • Embedded skills & wordlists: playbooks and wordlists baked into the binary, served on demand by tag. Skills are picked after fingerprinting the target, wordlists, large by nature, never cross the context window they're consumed server-side or iterated in-page.

Skills: embedded playbooks, dynamic selection

Playbooks live in the binary (go:embed over assets/skills/) and are served by two tools:

  • list_skills enumerates the available stacks (php, python, java, nodejs, ...).
  • load_skill returns the shared workflow with no arguments, or a stack's tailored playbook when you name it (load_skill(["python"])).

The flow: the model gets the shared workflow up front (so it knows its capabilities before it knows the target), fingerprints the target, then loads the matching stack. Polyglot targets are fine; call it again as more stacks surface. Adding a stack means dropping an assets/skills/<name>/ directory and rebuilding.

Quick start

go build -o mulot ./cmd/mulot        # build the MCP server binary

# local llama.cpp
python3 agent.py  --provider llamacpp --model qwen3.6-mtp --base http://localhost:8080 \
  "pwn this server http://localhost:4280/login.php and output only id and uname -a command"

# zai token
export ZAI_API_KEY=...        
python3 agent.py --provider zai --model glm-5.2 "audit http://localhost:8000"

agent.py is a small single-file harness that drives mulot over stdio MCP and talks to any OpenAI-compatible endpoint with tool calling. Presets: openrouter, llamacpp, zai. The provider is auto-detected from whichever API key is in the environment.

Environment variables:

VariableEffect
MULOT_USER_AGENTUser-Agent sent by http_request/http_fuzz and the browser.
MULOT_HEADLESSDefault headless mode for browser_launch (true/false).
MULOT_PROXYUpstream proxy for the browser (HTTP/SOCKS5) and for http_request/http_fuzz (HTTP/HTTPS only, SOCKS5 needs the browser tools).

Results

Running open-weights GLM-5.2 agents against web pentest challenges using only this harness (max 120 steps) produced strong results:

OverTheWire (Natas): 87% solved

Full results: 34 challenges