
pentest-ai-agents
Turn Claude Code into your offensive security research assistant. Specialized AI subagents for authorized penetration testing plan engagements,…

Turn Claude Code into your offensive security research assistant. Specialized AI subagents for authorized penetration testing plan engagements,…

An AI-powered threat modeling tool that leverages OpenAI's GPT models to generate threat models for a given application based on the STRIDE…

A modular framework for benchmarking LLMs and agentic strategies on security challenges across HackTheBox, TryHackMe, PortSwigger Labs, Cybench,…

Red Team AI Benchmark: Evaluating LLMs for authorized offensive-security tasks. Red Team AI Benchmark is a CLI model-evaluation benchmark. It…

Hands-on AI security lab platform with 50+ scenarios across prompt injection, agentic system exploitation, model manipulation, and MCP trust boundary…

Multi-agent automated context management for long horizon tasks in local AI Agents

Security research lab — Unauthenticated RCE via insecure deserialization in ComfyUI v0.23.0 (CVSS 9.8). Isolated Docker environment, technical…

The community's most comprehensive, continuously-updated index of research on Large Language Models for software vulnerability detection — papers…

Autonomous AI pentesting engine, continuous offensive security across web, cloud, AD & Kubernetes. Agentic reasoning + real exploit execution deliver…

Docker Compose lab reproducing CVE-2026-33626 SSRF in LMDeploy's vision-language image loader. Compares vulnerable (0.12.0) and patched (0.12.3)…

eBPF-powered network observability for Kubernetes. Indexes L4/L7 traffic with full K8s context, decrypts TLS without keys. Queryable by AI agents via…

Free XP on bug bounty, vulnerability scanning by wrapping well maintained tools, to perform automated tests. Alongside AI agents for binary analysis,…

Benchmark for evaluating safety risks of computer-using agents, with 104 realistic misuse scenarios across seven malicious categories, supporting…

Multi-agent static application-security review harness for AI coding agents: maps codebases, hunts vulnerability classes, chains and verifies…

Autonomous AI red team agent for penetration testing with 13+ specialized agents, 120+ OWASP test cases, and MITRE ATT&CK integration. Supports 15+…

AI red-team platform. Autonomous LLM agents run a penetration test end to end inside a Kali container and write the report. LangGraph plan/act…

Autonomous Hacking Agent for Red Team

A collection of challenge based hack-a-thons including student guide, coach guide, lecture presentations, sample/instructional code and templates. …