
The world's fastest agentic crawler. Reclaimed. Reinvented. Ready for war.
Agentic web crawler with anti-detection, vision fallback, and peer-to-peer mesh.
Endpoints · Mesh · Anti-Detection · Ghost Protocol · Live Stream · MCP Tools · Quick Start · Benchmarks · Architecture
Full-stack web crawling engine with JavaScript rendering, Camoufox anti-detect browser, per-request proxy routing, and autonomous agent loops. Converts pages to clean markdown with a native Rust extraction engine. When standard crawling is blocked by Cloudflare, CAPTCHAs, or JavaScript walls, Ghost Protocol captures a screenshot and extracts content via vision AI (Claude, GPT-4o, or Ollama). Supports multi-provider LLM orchestration across OpenAI, Anthropic, and Ollama in a single session. Nodes coordinate over a gossip-based peer-to-peer mesh for distributed crawling.
We integrated features from every major crawler — then added what none of them have.
| Feature | Crawl4AI | Firecrawl | Scrapy | Grub |
|---|---|---|---|---|
| JS rendering | ✅ Playwright | ✅ Playwright | ❌ HTTP only | ✅ Playwright |
| Anti-detect browser | stealth plugin | ❌ | ❌ | ✅ Camoufox |
| Ghost Protocol | ❌ | ❌ | ❌ | ✅ auto fallback |
| Per-request proxy | ✅ escalation | ❌ | middleware | ✅ per-request |
| Stealth patches | ✅ | ❌ | ❌ | ✅ opt-in |
| Agent loop | ✅ agentic | ✅ /agent | ❌ spiders | ✅ bounded SM |
| Live browser stream | ✅ WebSocket | ✅ Live View | ❌ | ✅ WS + MJPEG |
| Markdown output | ✅ Fit Markdown | ✅ core | ❌ | ✅ Rust engine |
| PDF extraction | ✅ PDF strategy | ✅ parse | ❌ | ✅ text layer + OCR fallback |
| MCP tools | ✅ community | ✅ official | ⚠️ community | ✅ 15 tools |
| Multi-provider LLM | ✅ all LLMs | ⚠️ Gemini | ❌ | ✅ OpenAI/Anthropic/Ollama |
| Mesh P2P | ❌ | ❌ | ❌ | ✅ gossip protocol |
| Policy enforcement | ❌ | ❌ | ❌ | ✅ domain gates + redaction |
| Prompt injection defense | ❌ | ❌ | ❌ | ✅ quarantine + visible-text diff |
| License | Apache 2.0 | AGPL-3.0 | BSD | BSD-3-Clause |
| Pricing | Free | Free–$333/mo | Free | Self-hosted |
| Feature | Browserbase | Scrapfly | Firecrawl Cloud | Grub |
|---|---|---|---|---|
| JS rendering | ✅ custom Chromium | ✅ proprietary | ✅ Playwright | ✅ Playwright |
| Anti-detect browser | ✅ custom Chromium | ✅ proprietary | ✅ cloud stealth | ✅ Camoufox |
| Ghost Protocol | ❌ | ❌ | ❌ | ✅ auto fallback |
| Per-request proxy | ✅ managed | ✅ 130M+ IPs | ✅ cloud-managed | ✅ per-request |
| Stealth patches | ✅ built-in | ✅ built-in | ✅ built-in | ✅ opt-in |
| Agent loop | ✅ Stagehand | ⚠️ via integrations | ✅ /agent | ✅ bounded SM |
| Live browser stream | ✅ iFrame + CDP | ✅ CDP | ✅ Live View | ✅ WS + MJPEG |
| Markdown output | ✅ via MCP | ✅ built-in | ✅ core | ✅ Rust engine |
| MCP tools | ✅ official | ✅ official | ✅ official | ✅ 15 tools |
| Mesh P2P | ❌ | ❌ | ❌ | ✅ gossip protocol |
| Policy enforcement | ❌ | ❌ | ❌ | ✅ domain gates + redaction |
| Prompt injection defense | ❌ | ❌ | ❌ | ✅ quarantine + visible-text diff |
| Self-hostable | ❌ cloud only | ❌ cloud only | ⚠️ limited OSS | ✅ full + Cloud Run |
| Pricing | Free–$99/mo | Usage-based | Free–$333/mo | Self-hosted |
Only Grub has Ghost Protocol — automatic vision-based fallback that screenshots blocked pages and extracts content via LLM when standard crawling fails. Prevention (Camoufox + proxy + stealth) handles 95% of blocks. Ghost Protocol handles the rest.
| Method | Path | Description | Status |
|---|---|---|---|
POST | /api/crawl | Single URL crawl (HTML + markdown) | Live |
POST | /api/markdown | Single or multi-URL markdown extraction | Live |
POST | /api/batch | Batch crawl with job tracking | Live |
POST | /api/raw | Raw HTML extraction (no markdown) | Live |
GET | /view | Browser-rendered HTML viewer | Live |
GET | /download | File download (PDFs, etc.) through crawler | Live |
POST | /api/pdf/pages | PDF pages as text + rendered PNG (base64) | Live |
PDF URLs are handled by /api/crawl, /api/markdown and /api/batch without the browser: the text layer is
extracted per page (PyMuPDF) and image-only pages fall back to the configured vision provider for OCR
(local default: Ollama with benhaotang/Nanonets-OCR-s; set AGENT_GHOST_VISION_PROVIDER=anthropic or openai
with a key to use a hosted model instead). OCR'd pages are labelled source: "ocr" with the model name and get a
<!-- ocr: <model> --> marker under their heading, so transcriptions are never mistaken for the source text.
Output is markdown with one ## Page N section per page; render_mode reports pdf_text, pdf_vision,
pdf_mixed or pdf_empty.
| Method | Path | Description | Status |
|---|---|---|---|
POST | /api/agent/run | Submit task to autonomous agent loop | Live |
GET | /api/agent/status/{run_id} | Check agent run status / load trace | Live |
POST | /api/agent/ghost | Ghost Protocol: screenshot + vision extract | Live |
| Method | Path | Description | Status |
|---|---|---|---|
POST | /api/jobs/create | Generic job submission | Live |
POST | /api/jobs/crawl | Submit single URL crawl job | Live |
POST | /api/jobs/batch-crawl | Submit batch crawl job | Live |
POST | /api/jobs/markdown | Submit markdown-only job | Live |
POST | /api/jobs/process-job | Cloud Tasks worker endpoint | Live |
POST | /api/wraith | AI-driven crawl workflow | Placeholder |
| Method | Path | Description | Status |
|---|---|---|---|
POST | /api/cache/search | Fuzzy search cached content | Live |
GET | /api/cache/list | List cached document metadata | Live |
GET | /api/cache/doc/{doc_id} | Fetch one cached document | Live |
POST | /api/cache/upsert | Upsert cache entries | Live |
POST | /api/cache/prune | Prune cache entries by TTL/domain | Live |