🔐 Guardian

Guardian is an enterprise-grade AI-powered penetration testing automation framework that combines multiple AI providers (OpenAI GPT-4, Claude, Google Gemini, OpenRouter, Requesty) with battle-tested security tools to deliver intelligent, adaptive security assessments with comprehensive evidence capture.
Features • Installation • Quick Start • Documentation • Contributing
⚠️ Legal Disclaimer
Guardian is designed exclusively for authorized security testing and educational purposes.
- ✅ Legal Use: Authorized penetration testing, security research, educational environments
- ❌ Illegal Use: Unauthorized access, malicious activities, any form of cyber attack
You are fully responsible for ensuring you have explicit written permission before testing any system. Unauthorized access to computer systems is illegal under laws including the Computer Fraud and Abuse Act (CFAA), GDPR, and equivalent international legislation.
By using Guardian, you agree to use it only on systems you own or have explicit authorization to test.
✨ Features
🤖 Multi-Provider AI Intelligence
- 7 AI Providers Supported: OpenAI (GPT-4o), Anthropic (Claude), Google (Gemini), OpenRouter, Requesty, Ollama (local), OpenAI-compatible (vLLM, LM Studio, Together, Groq)
- Plugin Provider Contract: Third-party providers ship via
[project.entry-points."guardian.providers"] — no fork required
- Multi-Agent Architecture: Specialized AI agents (Planner, Tool Selector, Analyst, Reporter) plus debate triage roles (Red Advocate, Blue Advocate, Judge) and Visual Triage
- Multi-Agent Debate Triage: Three-role red/blue/judge debate on ambiguous findings — F1 ≥ single-agent baseline +5pp
- Vision-LLM Visual Triage: Headless screenshot capture + image-grounded analyst enrichment via gpt-4o / Claude 3.5+ / Gemini 1.5+
- RAG Knowledge Base: SQLite + FTS5 grounded retrieval over CVE / CWE / MITRE ATT&CK feeds — kills hallucinated CVE refs
- Judge Model Routing:
think_deeply swap-and-restore — big model thinks, small model judges, ~10x cost reduction
- Learned Tool Selection: Offline ranker trained on session telemetry; abstains when low-confidence and falls back to LLM selector
- Adaptive Testing: AI adjusts tactics based on discovered vulnerabilities and prior tool yields
- False Positive Filtering: Debate triage cuts noise; cheap path skips when fp_probability is decisive
50 Integrated Security Tools across 10 categories:
| Category | Tools |
|---|
| Network | nmap, masscan |
| Web Reconnaissance | httpx, whatweb, wafw00f, cmseek |
| Subdomain / DNS | subfinder, amass, dnsrecon |
| Vulnerability Scanning | nuclei, nikto, sqlmap, wpscan |
| SSL/TLS Testing | testssl, sslyze |
| Content Discovery | gobuster, ffuf, arjun |
| Security Analysis | xsstrike, gitleaks |
| Cloud / Container / SBOM | trivy, grype, syft, scoutsuite, prowler, kube-bench |
| Modern Web + OSINT | graphw00f, clairvoyance, jwt_tool, shodan, theharvester |
| SAST + Secrets (B11) | semgrep, trufflehog, dependency-check |
| API Fuzzers (B10) | schemathesis, cariddi, restler |
| Burp/ZAP Bridge (B13) | zap, burp |
| LLM Red-Team (B12) | garak, pyrit, prompt_fuzz |
| Mobile Android (B9) | mobsf, apkleaks, objection |
| Active Directory (B8) | crackmapexec, bloodhound, kerbrute, impacket-secretsdump |
| Vision Evidence (A3) | playwright_screenshot |
📊 Enhanced Evidence Capture
- Execution Traceability: Every finding linked to its source tool execution via
execution_id
- Complete Command History: Full tool output preserved with each finding
- Raw Evidence Storage: Output snippets bound to findings
- Visual Evidence: Screenshots captured per URL, attached to web findings
- Session Reconstruction: Atomic-checkpointed
session_<id>.json enables --resume
🔄 Smart Workflow System (DSL v2)
- DAG Scheduler: Steps with
depends_on run in parallel up to max_parallel_tools
- Jinja2 Templates (sandboxed):
parameters: {key: "{{ <id>.parsed.alive_hosts }}"} resolves against prior step results
- Conditional Steps:
when: clauses gate execution on prior output
- Resume:
--resume picks up after the last completed step
- Parameter Priority: Workflow YAML > config block > tool defaults
- Custom Agents:
agent: debate | visual | analyst on analysis steps
- Multiple Report Formats: Markdown, HTML, JSON
📤 Output Integrations (B14)
- SARIF v2.1.0: GitHub-friendly, includes
security-severity, dedup fingerprints from execution_id
- DefectDojo: Direct REST upload
- Slack: Webhook posts with severity colour-coding
- Triggered via:
guardian report --export sarif --export defectdojo --export slack
🔒 Security & Compliance
- DNS-Resolve Scope Validation: Closes SSRF-class bypass; private RFC1918 ranges blacklisted
- Prompt-Injection Defense: All tool output wrapped via
<UNTRUSTED_TOOL_OUTPUT> delimiters + ANSI strip
- API Key Scrubbing: Logs and reports redact secrets at write time
- Confirmation Gate: Active+ tools (intrusive/destructive) require explicit user approval
- Audit Logging: Rotating logs of every AI decision and action
- Safe Mode: Prevents destructive actions by default
📋 Professional Reporting
- CVSS v3.1 Recomputation: Validates claimed scores against vector math; flags drift
- Executive Summaries: Non-technical overviews
- Technical Deep-Dives: Findings with evidence, CVSS, CWE, CVE, MITRE technique
- AI Decision Traces: Token usage, cost, thinking-chain ledger per agent
- Visual Triage Sections: Image-grounded enrichment baked into descriptions
- Async Throughout: Tool exec via
asyncio subprocess; agents async
- Lazy Tool Loading: 50 tools registered, none imported until needed —
--help stays under 500ms
- Parallel DAG Execution: Independent steps run concurrently per generation
- Workflow Automation: 13+ shipped workflows (recon, web, network, AD, mobile, LLM red-team, SAST, API)
📋 Prerequisites
Required
- Python 3.11 or higher (Download)
- AI Provider API Key (Choose one):
- Git (for cloning repository)
Guardian can intelligently use these tools if installed: