Skip to content
KitploitKITPLOIT
ToolsExploitsBlog
Log in
Submit
ToolsExploitsBlog
Submit

Hacking, PenTest, and Cybersecurity Tools for Your Security Arsenal!

Kitploit is a directory of hacking, cybersecurity, and pentesting tools. Discover the latest project updates to find vulnerabilities, analyze systems, automate testing, and strengthen your security.

··Feeds·Contact·Privacy·© 2026 Kitploit

Tool Directory

Categories

View all categories
Loading categories
grubcrawler — The world's fastest agentic crawler. Reclaimed. Reinvented. Ready for war. | Kitploit
Tools/GitHubGitHub/deepbluedynamics/grubcrawler
OSINT (Open Source Intelligence)ReconnaissanceDynamic Analysis (Sandboxing)Web Proxies & InterceptionInformation GatheringWeb SecurityPenetration TestingUtilities & FrameworksMachine LearningRed TeamingCrawlerAnti-Bot
347191 day agoReviewed by Kitploit

Most Popular

View all →

Discover the most used tools by our community.

Explore all tools

Browse our collection of tools

View all tools →
Share
GitHubdeepbluedynamics/grubcrawler

grubcrawler

The world's fastest agentic crawler. Reclaimed. Reinvented. Ready for war.

View RepositoryWebsite
Grub Crawler

License Python FastAPI Playwright MCP Ghost Protocol Live Stream Camoufox Proxy


Agentic web crawler with anti-detection, vision fallback, and peer-to-peer mesh.


Endpoints · Mesh · Anti-Detection · Ghost Protocol · Live Stream · MCP Tools · Quick Start · Benchmarks · Architecture


Full-stack web crawling engine with JavaScript rendering, Camoufox anti-detect browser, per-request proxy routing, and autonomous agent loops. Converts pages to clean markdown with a native Rust extraction engine. When standard crawling is blocked by Cloudflare, CAPTCHAs, or JavaScript walls, Ghost Protocol captures a screenshot and extracts content via vision AI (Claude, GPT-4o, or Ollama). Supports multi-provider LLM orchestration across OpenAI, Anthropic, and Ollama in a single session. Nodes coordinate over a gossip-based peer-to-peer mesh for distributed crawling.


Why Grub

We integrated features from every major crawler — then added what none of them have.

Self-Hosted Crawlers

FeatureCrawl4AIFirecrawlScrapyGrub
JS rendering✅ Playwright✅ Playwright❌ HTTP only✅ Playwright
Anti-detect browserstealth plugin❌❌✅ Camoufox
Ghost Protocol❌❌❌✅ auto fallback
Per-request proxy✅ escalation❌middleware✅ per-request
Stealth patches✅❌❌✅ opt-in
Agent loop✅ agentic✅ /agent❌ spiders✅ bounded SM
Live browser stream✅ WebSocket✅ Live View❌✅ WS + MJPEG
Markdown output✅ Fit Markdown✅ core❌✅ Rust engine
PDF extraction✅ PDF strategy✅ parse❌✅ text layer + OCR fallback
MCP tools✅ community✅ official⚠️ community✅ 15 tools
Multi-provider LLM✅ all LLMs⚠️ Gemini❌✅ OpenAI/Anthropic/Ollama
Mesh P2P❌❌❌✅ gossip protocol
Policy enforcement❌❌❌✅ domain gates + redaction
Prompt injection defense❌❌❌✅ quarantine + visible-text diff
LicenseApache 2.0AGPL-3.0BSDBSD-3-Clause
PricingFreeFree–$333/moFreeSelf-hosted

Cloud / Managed Crawlers

FeatureBrowserbaseScrapflyFirecrawl CloudGrub
JS rendering✅ custom Chromium✅ proprietary✅ Playwright✅ Playwright
Anti-detect browser✅ custom Chromium✅ proprietary✅ cloud stealth✅ Camoufox
Ghost Protocol❌❌❌✅ auto fallback
Per-request proxy✅ managed✅ 130M+ IPs✅ cloud-managed✅ per-request
Stealth patches✅ built-in✅ built-in✅ built-in✅ opt-in
Agent loop✅ Stagehand⚠️ via integrations✅ /agent✅ bounded SM
Live browser stream✅ iFrame + CDP✅ CDP✅ Live View✅ WS + MJPEG
Markdown output✅ via MCP✅ built-in✅ core✅ Rust engine
MCP tools✅ official✅ official✅ official✅ 15 tools
Mesh P2P❌❌❌✅ gossip protocol
Policy enforcement❌❌❌✅ domain gates + redaction
Prompt injection defense❌❌❌✅ quarantine + visible-text diff
Self-hostable❌ cloud only❌ cloud only⚠️ limited OSS✅ full + Cloud Run
PricingFree–$99/moUsage-basedFree–$333/moSelf-hosted

Only Grub has Ghost Protocol — automatic vision-based fallback that screenshots blocked pages and extracts content via LLM when standard crawling fails. Prevention (Camoufox + proxy + stealth) handles 95% of blocks. Ghost Protocol handles the rest.

API Endpoints

Core Crawling

MethodPathDescriptionStatus
POST/api/crawlSingle URL crawl (HTML + markdown)Live
POST/api/markdownSingle or multi-URL markdown extractionLive
POST/api/batchBatch crawl with job trackingLive
POST/api/rawRaw HTML extraction (no markdown)Live
GET/viewBrowser-rendered HTML viewerLive
GET/downloadFile download (PDFs, etc.) through crawlerLive
POST/api/pdf/pagesPDF pages as text + rendered PNG (base64)Live

PDF URLs are handled by /api/crawl, /api/markdown and /api/batch without the browser: the text layer is extracted per page (PyMuPDF) and image-only pages fall back to the configured vision provider for OCR (local default: Ollama with benhaotang/Nanonets-OCR-s; set AGENT_GHOST_VISION_PROVIDER=anthropic or openai with a key to use a hosted model instead). OCR'd pages are labelled source: "ocr" with the model name and get a <!-- ocr: <model> --> marker under their heading, so transcriptions are never mistaken for the source text. Output is markdown with one ## Page N section per page; render_mode reports pdf_text, pdf_vision, pdf_mixed or pdf_empty.

Agent (Mode B)

MethodPathDescriptionStatus
POST/api/agent/runSubmit task to autonomous agent loopLive
GET/api/agent/status/{run_id}Check agent run status / load traceLive
POST/api/agent/ghostGhost Protocol: screenshot + vision extractLive

Job Management

MethodPathDescriptionStatus
POST/api/jobs/createGeneric job submissionLive
POST/api/jobs/crawlSubmit single URL crawl jobLive
POST/api/jobs/batch-crawlSubmit batch crawl jobLive
POST/api/jobs/markdownSubmit markdown-only jobLive
POST/api/jobs/process-jobCloud Tasks worker endpointLive
POST/api/wraithAI-driven crawl workflowPlaceholder

Remote Cache

MethodPathDescriptionStatus
POST/api/cache/searchFuzzy search cached contentLive
GET/api/cache/listList cached document metadataLive
GET/api/cache/doc/{doc_id}Fetch one cached documentLive
POST/api/cache/upsertUpsert cache entriesLive
POST/api/cache/prunePrune cache entries by TTL/domainLive
Download Tool