Autonomous Hacking Agent for Red Team
"Another AI hacker? Let us guess — it runs nmap and writes a report."
Skip the Docker setup — run autonomous red-team engagements right from your browser.
Prerequisites: Docker and Docker Compose v2. Supported on macOS (Apple Silicon + Intel), Linux (amd64 + arm64), and Windows (amd64 + arm64) — native via PowerShell or via WSL2 (Ubuntu / Kali).
macOS / Linux / WSL2
curl -fsSL https://decepticon.red/install | bash
decepticon onboard # Interactive setup wizard (provider, API key, model profile)
decepticon # Start the core stack and drop into the terminal CLI
The default start brings up the core management plane (LiteLLM, PostgreSQL, Neo4j, Skillogy, LangGraph, sandbox) and launches the terminal CLI. Specialist workloads (BloodHound CE, Sliver C2, Ghidra MCP, …) and the web dashboard come up on demand — the orchestrator spawns specialists via ops_start("ad") etc., and you bring up the dashboard from inside the CLI with /web (see Web Dashboard).
Windows (PowerShell, native)
irm https://decepticon.red/install.ps1 | iex
decepticon onboard
decepticon
→ Quick start · Full setup walkthrough
Building on top of the agents — a product, a research integration, or a custom orchestrator? Install the SDK from PyPI:
pip install decepticon # core SDK
pip install "decepticon[neo4j]" # + the knowledge-graph attack-chain tools
decepticon is a client SDK: it ships the agent factories, middleware, tools, and skills, and routes LLM calls and sandbox execution to runtime services over HTTP (DECEPTICON_LLM__PROXY_URL, SANDBOX_URL). Running agents still needs those services — use the Docker stack above, or point the URLs at your own equivalents. See Decepticon as a library for the factory override surface, declarative PluginBundle plugins, and the safety gate.
We're building Decepticon toward an Offensive Vaccine for the AI-driven threat landscape. If you believe in autonomous red teaming as a path to stronger defense, consider supporting the project.
| Benchmark | Difficulty | Pass Rate |
|---|---|---|
| XBOW validation-benchmarks | Easy (Level 1) | 45 / 45 (100 %) |
| XBOW validation-benchmarks | Medium (Level 2) | 50 / 51 (98.0 %) |
| XBOW validation-benchmarks | Hard (Level 3) | 7 / 8 (87.5 %) |
| XBOW validation-benchmarks | All levels | 102 / 104 (98.08 %) |
The "AI + hacking" space is full of demos that run nmap and print a report. That's not what this is.
Decepticon is a professional autonomous Red Team agent. It executes realistic attack chains — reconnaissance, exploitation, privilege escalation, lateral movement, C2 — the way a real adversary would, not the way a scanner does.
But more importantly: it operates under the discipline that separates red teamers from script kiddies. Before a single packet leaves the wire, Decepticon generates a complete engagement package — RoE, ConOps, Deconfliction Plan, and OPPLAN with MITRE ATT&CK mapping — and every action runs inside those defined rules.
→ Engagement workflow deep dive
Real kill chains, not checkbox scans. Decepticon reads an OPPLAN and pursues objectives through whatever path opens up — pivoting, adapting, chaining techniques.
Interactive shells, actually. Real offensive tools are interactive (msfconsole, sliver-client, evil-winrm). Decepticon runs every command inside persistent tmux sessions with automatic prompt detection — so when a tool drops into an interactive prompt, the agent sends follow-up commands without workarounds.
Hardened sandbox isolation. All commands run inside a Kali Linux sandbox on a dedicated operational network (sandbox-net), separate from the management plane (decepticon-net). LangGraph drives the sandbox via the Docker socket. → Architecture