
Fully autonomous AI Agents system capable of performing complex penetration testing tasks
Join the Community! Connect with security researchers, AI enthusiasts, and fellow ethical hackers. Get support, share insights, and stay updated with the latest PentAGI developments.
PentAGI is an innovative tool for automated security testing that leverages cutting-edge artificial intelligence technologies. The project is designed for information security professionals, researchers, and enthusiasts who need a powerful and flexible solution for conducting penetration tests.
flowchart TB
classDef person fill:#08427B,stroke:#073B6F,color:#fff
classDef system fill:#1168BD,stroke:#0B4884,color:#fff
classDef external fill:#666666,stroke:#0B4884,color:#fff
pentester["👤 Security Engineer
(User of the system)"]
pentagi["✨ PentAGI
(Autonomous penetration testing system)"]
target["🎯 target-system
(System under test)"]
llm["🧠 llm-provider
(OpenAI/Anthropic/Ollama/Bedrock/Gemini/Custom)"]
search["🔍 search-systems
(Google/DuckDuckGo/Tavily/Traversaal/Perplexity/Sploitus/Searxng)"]
langfuse["📊 langfuse-ui
(LLM Observability Dashboard)"]
grafana["📈 grafana
(System Monitoring Dashboard)"]
pentester --> |Uses HTTPS| pentagi
pentester --> |Monitors AI HTTPS| langfuse
pentester --> |Monitors System HTTPS| grafana
pentagi --> |Tests Various protocols| target
pentagi --> |Queries HTTPS| llm
pentagi --> |Searches HTTPS| search
pentagi --> |Reports HTTPS| langfuse
pentagi --> |Reports HTTPS| grafana
class pentester person
class pentagi system
class target,llm,search,langfuse,grafana external
linkStyle default stroke:#ffffff,color:#ffffffgraph TB
subgraph Core Services
UI[Frontend UI<br/>React + TypeScript]
API[Backend API<br/>Go + GraphQL]
DB[(Vector Store<br/>PostgreSQL + pgvector)]
MQ[Task Queue<br/>Async Processing]
Agent[AI Agents<br/>Multi-Agent System]
end
subgraph Knowledge Graph
Graphiti[Graphiti<br/>Knowledge Graph API]
Neo4j[(Neo4j<br/>Graph Database)]
end
subgraph Monitoring
Grafana[Grafana<br/>Dashboards]
VictoriaMetrics[VictoriaMetrics<br/>Time-series DB]
Jaeger[Jaeger<br/>Distributed Tracing]
Loki[Loki<br/>Log Aggregation]
OTEL[OpenTelemetry<br/>Data Collection]
end
subgraph Analytics
Langfuse[Langfuse<br/>LLM Analytics]
ClickHouse[ClickHouse<br/>Analytics DB]
Redis[Redis<br/>Cache + Rate Limiter]
MinIO[MinIO<br/>S3 Storage]
end
subgraph Security Tools
Scraper[Web Scraper<br/>Isolated Browser]
PenTest[Security Tools<br/>20+ Pro Tools<br/>Sandboxed Execution]
end
UI --> |HTTP/WS| API
API --> |SQL| DB
API --> |Events| MQ
MQ --> |Tasks| Agent
Agent --> |Commands| PenTest
Agent --> |Queries| DB
Agent --> |Knowledge| Graphiti
Graphiti --> |Graph| Neo4j
API --> |Telemetry| OTEL
OTEL --> |Metrics| VictoriaMetrics
OTEL --> |Traces| Jaeger
OTEL --> |Logs| Loki
Grafana --> |Query| VictoriaMetrics
Grafana --> |Query| Jaeger
Grafana --> |Query| Loki
API --> |Analytics| Langfuse
Langfuse --> |Store| ClickHouse
Langfuse --> |Cache| Redis
Langfuse --> |Files| MinIO
classDef core fill:#f9f,stroke:#333,stroke-width:2px,color:#000
classDef knowledge fill:#ffa,stroke:#333,stroke-width:2px,color:#000
classDef monitoring fill:#bbf,stroke:#333,stroke-width:2px,color:#000
classDef analytics fill:#bfb,stroke:#333,stroke-width:2px,color:#000
classDef tools fill:#fbb,stroke:#333,stroke-width:2px,color:#000
class UI,API,DB,MQ,Agent core
class Graphiti,Neo4j knowledge
class Grafana,VictoriaMetrics,Jaeger,Loki,OTEL monitoring
class Langfuse,ClickHouse,Redis,MinIO analytics
class Scraper,PenTest toolserDiagram
Flow ||--o{ Task : contains
Task ||--o{ SubTask : contains
SubTask ||--o{ Action : contains
Action ||--o{ Artifact : produces
Action ||--o{ Memory : stores
Flow {
string id PK
string name "Flow name"
string description "Flow description"
string status "active/completed/failed"
json parameters "Flow parameters"
timestamp created_at
timestamp updated_at
}
Task {
string id PK
string flow_id FK
string name "Task name"
string description "Task description"
string status "pending/running/done/failed"
json result "Task results"
timestamp created_at
timestamp updated_at
}
SubTask {
string id PK
string task_id FK
string name "Subtask name"
string description "Subtask description"
string status "queued/running/completed/failed"
string agent_type "researcher/developer/executor"
json context "Agent context"
timestamp created_at
timestamp updated_at
}
Action {
string id PK
string subtask_id FK
string type "command/search/analyze/etc"
string status "success/failure"
json parameters "Action parameters"
json result "Action results"
timestamp created_at
}
Artifact {
string id PK
string action_id FK
string type "file/report/log"
string path "Storage path"
json metadata "Additional info"
timestamp created_at
}
Memory {
string id PK
string action_id FK
string type "observation/conclusion"
vector embedding "Vector representation"
text content "Memory content"
timestamp created_at
}sequenceDiagram
participant O as Orchestrator
participant R as Researcher
participant D as Developer
participant E as Executor
participant VS as Vector Store
participant KB as Knowledge Base
Note over O,KB: Flow Initialization
O->>VS: Query similar tasks
VS-->>O: Return experiences
O->>KB: Load relevant knowledge
KB-->>O: Return context
Note over O,R: Research Phase
O->>R: Analyze target
R->>VS: Search similar cases
VS-->>R: Return patterns
R->>KB: Query vulnerabilities
KB-->>R: Return known issues
R->>VS: Store findings
R-->>O: Research results
Note over O,D: Planning Phase
O->>D: Plan attack
D->>VS: Query exploits
VS-->>D: Return techniques
D->>KB: Load tools info
KB-->>D: Return capabilities
D-->>O: Attack plan
Note over O,E: Execution Phase
O->>E: Execute plan
E->>KB: Load tool guides
KB-->>E: Return procedures
E->>VS: Store results
E-->>O: Execution statusgraph TB
subgraph "Long-term Memory"
VS[(Vector Store<br/>Embeddings DB)]
KB[Knowledge Base<br/>Domain Expertise]
Tools[Tools Knowledge<br/>Usage Patterns]
end
subgraph "Working Memory"
Context[Current Context<br/>Task State]
Goals[Active Goals<br/>Objectives]
State[System State<br/>Resources]
end
subgraph "Episodic Memory"
Actions[Past Actions<br/>Commands History]
Results[Action Results<br/>Outcomes]
Patterns[Success Patterns<br/>Best Practices]
end
Context --> |Query| VS
VS --> |Retrieve| Context
Goals --> |Consult| KB
KB --> |Guide| Goals
State --> |Record| Actions
Actions --> |Learn| Patterns
Patterns --> |Store| VS
Tools --> |Inform| State
Results --> |Update| Tools
VS --> |Enhance| KB
KB --> |Index| VS
classDef ltm fill:#f9f,stroke:#333,stroke-width:2px,color:#000
classDef wm fill:#bbf,stroke:#333,stroke-width:2px,color:#000
classDef em fill:#bfb,stroke:#333,stroke-width:2px,color:#000
class VS,KB,Tools ltm
class Context,Goals,State wm
class Actions,Results,Patterns emThe chain summarization system manages conversation context growth by selectively summarizing older messages. This is critical for preventing token limits from being exceeded while maintaining conversation coherence.
flowchart TD
A[Input Chain] --> B{Needs Summarization?}
B -->|No| C[Return Original Chain]
B -->|Yes| D[Convert to ChainAST]
D --> E[Apply Section Summarization]
E --> F[Process Oversized Pairs]
F --> G[Manage Last Section Size]
G --> H[Apply QA Summarization]
H --> I[Rebuild Chain with Summaries]
I --> J{Is New Chain Smaller?}
J -->|Yes| K[Return Optimized Chain]
J -->|No| C
classDef process fill:#bbf,stroke:#333,stroke-width:2px,color:#000
classDef decision fill:#bfb,stroke:#333,stroke-width:2px,color:#000
classDef output fill:#fbb,stroke:#333,stroke-width:2px,color:#000
class A,D,E,F,G,H,I process
class B,J decision
class C,K outputThe algorithm operates on a structured representation of conversation chains (ChainAST) that preserves message types including tool calls and their responses. All summarization operations maintain critical conversation flow while reducing context size.
| Parameter | Environment Variable | Default | Description |
|---|---|---|---|
| Preserve Last | SUMMARIZER_PRESERVE_LAST | true | Whether to keep all messages in the last section intact |
| Use QA Pairs | SUMMARIZER_USE_QA | true | Whether to use QA pair summarization strategy |
| Summarize Human in QA | SUMMARIZER_SUM_MSG_HUMAN_IN_QA | false | Whether to summarize human messages in QA pairs |
| Last Section Size | SUMMARIZER_LAST_SEC_BYTES | 51200 | Maximum byte size for last section (50KB) |
| Max Body Pair Size | SUMMARIZER_MAX_BP_BYTES | 16384 | Maximum byte size for a single body pair (16KB) |
| Max QA Sections | SUMMARIZER_MAX_QA_SECTIONS | 10 | Maximum QA pair sections to preserve |
| Max QA Size | SUMMARIZER_MAX_QA_BYTES | 65536 | Maximum byte size for QA pair sections (64KB) |
| Keep QA Sections | SUMMARIZER_KEEP_QA_SECTIONS | 1 | Number of recent QA sections to keep without summarization |
Assistant instances can use customized summarization settings to fine-tune context management behavior:
The assistant summarizer configuration provides more memory for context retention compared to the global settings, preserving more recent conversation history while still ensuring efficient token usage.
# Default values for global summarizer logic
SUMMARIZER_PRESERVE_LAST=true
SUMMARIZER_USE_QA=true
SUMMARIZER_SUM_MSG_HUMAN_IN_QA=false
SUMMARIZER_LAST_SEC_BYTES=51200
SUMMARIZER_MAX_BP_BYTES=16384
SUMMARIZER_MAX_QA_SECTIONS=10
SUMMARIZER_MAX_QA_BYTES=65536
SUMMARIZER_KEEP_QA_SECTIONS=1
# Default values for assistant summarizer logic
ASSISTANT_SUMMARIZER_PRESERVE_LAST=true
ASSISTANT_SUMMARIZER_LAST_SEC_BYTES=76800
ASSISTANT_SUMMARIZER_MAX_BP_BYTES=16384
ASSISTANT_SUMMARIZER_MAX_QA_SECTIONS=7
ASSISTANT_SUMMARIZER_MAX_QA_BYTES=76800
ASSISTANT_SUMMARIZER_KEEP_QA_SECTIONS=3
PentAGI includes sophisticated multi-layered agent supervision mechanisms to ensure efficient task execution, prevent infinite loops, and provide intelligent recovery from stuck states:
<original_result> and <mentor_analysis> sectionsEXECUTION_MONITOR_ENABLED (default: false), customize thresholds with EXECUTION_MONITOR_SAME_TOOL_LIMIT and EXECUTION_MONITOR_TOTAL_TOOL_LIMITBest for: Smaller models (< 32B parameters), complex attack scenarios requiring continuous guidance, preventing agents from getting stuck on single approach
Performance Impact: 2-3x increase in execution time and token usage, but delivers 2x improvement in result quality based on testing with Qwen3.5-27B-FP8
<task_assignment> structure with execution plan and instructionsAGENT_PLANNING_STEP_ENABLED (default: false)Best for: Models < 32B parameters, complex penetration testing workflows, improving success rates on sophisticated tasks
Enhanced Adviser Configuration: Works exceptionally well when adviser agent uses stronger model or enhanced settings. Example: using same base model with maximum reasoning mode for adviser (see vllm-qwen3.5-27b-fp8.provider.yml) enables comprehensive task analysis and strategic planning from identical model architecture.
Performance Impact: Adds planning overhead but significantly improves completion rates and reduces redundant work
MAX_GENERAL_AGENT_TOOL_CALLS (default: 100)MAX_LIMITED_AGENT_TOOL_CALLS (default: 20)done, ask)Must-Have for Models < 32B Parameters: Testing with Qwen3.5-27B-FP8 demonstrates that enabling both Execution Monitoring and Task Planning is essential for smaller open source models:
Trade-offs:
Configuration Strategy: For optimal performance with smaller models, configure adviser agent with enhanced settings:
vllm-qwen3.5-27b-fp8.provider.yml)The architecture of PentAGI is designed to be modular, scalable, and secure. Here are the key components:
Core Services
Knowledge Graph
Monitoring Stack
Analytics Platform
Security Tools
Memory Systems
The system uses Docker containers for isolation and easy deployment, with separate networks for core services, monitoring, and analytics to ensure proper security boundaries. Each component is designed to scale horizontally and can be configured for high availability in production environments.
PentAGI provides an interactive installer with a terminal-based UI for streamlined configuration and deployment. The installer guides you through system checks, LLM provider setup, search engine configuration, and security hardening.
Supported Platforms:
Quick Installation (Linux amd64):
# Create installation directory
mkdir -p pentagi && cd pentagi
# Download installer
wget -O installer.zip https://pentagi.com/downloads/linux/amd64/installer-latest.zip
# Extract
unzip installer.zip
# Run interactive installer
./installer
Prerequisites & Permissions:
The installer requires appropriate privileges to interact with the Docker API for proper operation. By default, it uses the Docker socket (/var/run/docker.sock) which requires either:
Option 1 (Recommended for production): Run the installer as root:
sudo ./installer
Option 2 (Development environments): Grant your user access to the Docker socket by adding them to the docker group:
# Add your user to the docker group
sudo usermod -aG docker $USER
# Log out and log back in, or activate the group immediately
newgrp docker
# Verify Docker access (should run without sudo)
docker ps
⚠️ Security Note: Adding a user to the docker group grants root-equivalent privileges. Only do this for trusted users in controlled environments. For production deployments, consider using rootless Docker mode or running the installer with sudo.
The installer will:
.env file with optimal defaultsThe PentAGI web console already manages several settings areas after the server is up and running:
The following configuration areas still need to be set on the server through environment variables, compose files, or mounted config files:
OLLAMA_SERVER_CONFIG_PATH and LLM_SERVER_CONFIG_PATH.DUCKDUCKGO_*, GOOGLE_*, TAVILY_API_KEY, TRAVERSAAL_API_KEY, PERPLEXITY_*, SEARXNG_*, and SPLOITUS_ENABLED.For Production & Enhanced Security:
For production deployments or security-sensitive environments, we strongly recommend using a distributed two-node architecture where worker operations are isolated on a separate server. This prevents untrusted code execution and network access issues on your main system.
See detailed guide: Worker Node Setup
The two-node setup provides:
mkdir pentagi && cd pentagi
.env.example to .env or download it:curl -o .env https://raw.githubusercontent.com/vxcontrol/pentagi/master/.env.example
example.custom.provider.yml, example.ollama.provider.yml) or download it:curl -o example.custom.provider.yml https://raw.githubusercontent.com/vxcontrol/pentagi/master/examples/configs/custom-openai.provider.yml
curl -o example.ollama.provider.yml https://raw.githubusercontent.com/vxcontrol/pentagi/master/examples/configs/ollama-llama318b.provider.yml
.env file.# Required: At least one of these LLM providers
OPEN_AI_KEY=your_openai_key
ANTHROPIC_API_KEY=your_anthropic_key
GEMINI_API_KEY=your_gemini_key
# Optional: AWS Bedrock provider (enterprise-grade models)
BEDROCK_REGION=us-east-1
# Choose one authentication method:
BEDROCK_DEFAULT_AUTH=true # Option 1: Use AWS SDK default credential chain (recommended for EC2/ECS)
# BEDROCK_BEARER_TOKEN=your_bearer_token # Option 2: Bearer token authentication
# BEDROCK_ACCESS_KEY_ID=your_aws_access_key # Option 3: Static credentials
# BEDROCK_SECRET_ACCESS_KEY=your_aws_secret_key
# Optional: Ollama provider (local or cloud)
# OLLAMA_SERVER_URL=http://ollama-server:11434 # Local server
# OLLAMA_SERVER_URL=https://ollama.com # Cloud service
# OLLAMA_SERVER_API_KEY=your_ollama_cloud_key # Required for cloud, empty for local
# Optional: Chinese AI providers
# DEEPSEEK_API_KEY=your_deepseek_key # DeepSeek (strong reasoning)
# GLM_API_KEY=your_glm_key # GLM (Zhipu AI)
# KIMI_API_KEY=your_kimi_key # Kimi (Moonshot AI, ultra-long context)
# QWEN_API_KEY=your_qwen_key # Qwen (Alibaba Cloud, multimodal)
# Optional: Local LLM provider (zero-cost inference)
OLLAMA_SERVER_URL=http://localhost:11434
OLLAMA_SERVER_MODEL=your_model_name
# Optional: Additional search capabilities
DUCKDUCKGO_ENABLED=true
DUCKDUCKGO_REGION=us-en
DUCKDUCKGO_SAFESEARCH=
DUCKDUCKGO_TIME_RANGE=
SPLOITUS_ENABLED=true
GOOGLE_API_KEY=your_google_key
GOOGLE_CX_KEY=your_google_cx
TAVILY_API_KEY=your_tavily_key
TRAVERSAAL_API_KEY=your_traversaal_key
PERPLEXITY_API_KEY=your_perplexity_key
PERPLEXITY_MODEL=sonar-pro
PERPLEXITY_CONTEXT_SIZE=medium
# Searxng meta search engine (aggregates results from multiple sources)
SEARXNG_URL=http://your-searxng-instance:8080
SEARXNG_CATEGORIES=general
SEARXNG_LANGUAGE=
SEARXNG_SAFESEARCH=0
SEARXNG_TIME_RANGE=
SEARXNG_TIMEOUT=
## Graphiti knowledge graph settings
GRAPHITI_ENABLED=true
GRAPHITI_TIMEOUT=30
GRAPHITI_URL=http://graphiti:8000
GRAPHITI_MODEL_NAME=gpt-5-mini
# Neo4j settings (used by Graphiti stack)
NEO4J_USER=neo4j
NEO4J_DATABASE=neo4j
NEO4J_PASSWORD=devpassword
NEO4J_URI=bolt://neo4j:7687
# Assistant configuration
ASSISTANT_USE_AGENTS=false # Default value for agent usage when creating new assistants
.env file to improve security.COOKIE_SIGNING_SALT - Salt for cookie signing, change to random valuePUBLIC_URL - Public URL of your server (eg. https://pentagi.example.com)SERVER_SSL_CRT and SERVER_SSL_KEY - Custom paths to your existing SSL certificate and key for HTTPS (these paths should be used in the docker-compose.yml file to mount as volumes)SCRAPER_PUBLIC_URL - Public URL for scraper if you want to use different scraper server for public URLsSCRAPER_PRIVATE_URL - Private URL for scraper (local scraper server in docker-compose.yml file to access it to local URLs)PENTAGI_POSTGRES_USER and PENTAGI_POSTGRES_PASSWORD - PostgreSQL credentialsNEO4J_USER and NEO4J_PASSWORD - Neo4j credentials (for Graphiti knowledge graph).env file if you want to use it in VSCode or other IDEs as a envFile option:perl -i -pe 's/\s+#.*$//' .env
curl -O https://raw.githubusercontent.com/vxcontrol/pentagi/master/docker-compose.yml
docker compose up -d
Visit localhost:8443 to access PentAGI Web UI (default is [email protected] / admin)
PentAGI does not expose public self-service sign-up from the login page. A fresh installation creates the default local administrator account:
[email protected]adminOn first login, change the default password before using the instance for real work. If the administrator password is lost later, use the installer maintenance menu to reset the default [email protected] account password.
For multi-user setups, an authenticated administrator can manage local users through the Users REST API (/api/v1/users/). The OpenAPI UI is available at https://localhost:8443/api/v1/swagger/index.html after the instance is running.
[!NOTE] If you caught an error about
pentagi-networkorobservability-networkorlangfuse-networkyou need to rundocker-compose.ymlfirstly to create these networks and after that rundocker-compose-langfuse.yml,docker-compose-graphiti.yml, anddocker-compose-observability.ymlto use Langfuse, Graphiti, and Observability services.You have to set at least one Language Model provider (OpenAI, Anthropic, Gemini, AWS Bedrock, or Ollama) to use PentAGI. AWS Bedrock provides enterprise-grade access to multiple foundation models from leading AI companies, while Ollama provides zero-cost local inference if you have sufficient computational resources. Additional API keys for search engines are optional but recommended for better results.
For fully local deployment with advanced models: See our comprehensive guide on Running PentAGI with vLLM and Qwen3.5-27B-FP8 for a production-grade local LLM setup. This configuration achieves ~13,000 TPS for prompt processing and ~650 TPS for completion on 4× RTX 5090 GPUs, supporting 12+ concurrent flows with complete independence from cloud providers.
LLM_SERVER_*environment variables are experimental feature and will be changed in the future. Right now you can use them to specify custom LLM server URL and one model for all agent types.
PROXY_URLis a global proxy URL for all LLM providers and external search systems. You can use it for isolation from external networks.The
docker-compose.ymlfile runs the PentAGI service as root user because it needs access to docker.sock for container management. If you're using TCP/IP network connection to Docker instead of socket file, you can remove root privileges and use the defaultpentagiuser for better security.
By default, PentAGI binds to 127.0.0.1 (localhost only) for security. To access PentAGI from other machines on your network, you need to configure external access.
.env file with your server's IP address:# Network binding - allow external connections
PENTAGI_LISTEN_IP=0.0.0.0
PENTAGI_LISTEN_PORT=8443
# Public URL - use your actual server IP or hostname
# Replace 192.168.1.100 with your server's IP address
PUBLIC_URL=https://192.168.1.100:8443
# CORS origins - list all URLs that will access PentAGI
# Include localhost for local access AND your server IP for external access
CORS_ORIGINS=https://localhost:8443,https://192.168.1.100:8443
[!IMPORTANT]
- Replace
192.168.1.100with your actual server's IP address- Do NOT use
0.0.0.0inPUBLIC_URLorCORS_ORIGINS- use the actual IP address- Include both localhost and your server IP in
CORS_ORIGINSfor flexibility
docker compose down
docker compose up -d --force-recreate
docker ps | grep pentagi
You should see 0.0.0.0:8443->8443/tcp or :::8443->8443/tcp.
If you see 127.0.0.1:8443->8443/tcp, the environment variable wasn't picked up. In this case, directly edit docker-compose.yml line 31:
ports:
- "0.0.0.0:8443:8443"
Then recreate containers again.
# Ubuntu/Debian with UFW
sudo ufw allow 8443/tcp
sudo ufw reload
# CentOS/RHEL with firewalld
sudo firewall-cmd --permanent --add-port=8443/tcp
sudo firewall-cmd --reload
https://localhost:8443https://your-server-ip:8443[!NOTE] You'll need to accept the self-signed SSL certificate warning in your browser when accessing via IP address.
PentAGI fully supports Podman as a Docker alternative. However, when using Podman in rootless mode, the scraper service requires special configuration because rootless containers cannot bind privileged ports (ports below 1024).
The default scraper configuration uses port 443 (HTTPS), which is a privileged port. For Podman rootless, reconfigure the scraper to use a non-privileged port:
1. Edit docker-compose.yml - modify the scraper service (around line 199):
scraper:
image: vxcontrol/scraper:latest
restart: unless-stopped
container_name: scraper
hostname: scraper
expose:
- 3000/tcp # Changed from 443 to 3000
ports:
- "${SCRAPER_LISTEN_IP:-127.0.0.1}:${SCRAPER_LISTEN_PORT:-9443}:3000" # Map to port 3000
environment:
- MAX_CONCURRENT_SESSIONS=${LOCAL_SCRAPER_MAX_CONCURRENT_SESSIONS:-10}
- USERNAME=${LOCAL_SCRAPER_USERNAME:-someuser}
- PASSWORD=${LOCAL_SCRAPER_PASSWORD:-somepass}
logging:
options:
max-size: 50m
max-file: "7"
volumes:
- scraper-ssl:/usr/src/app/ssl
networks:
- pentagi-network
shm_size: 2g
2. Update .env file - change the scraper URL to use HTTP and port 3000:
# Scraper configuration for Podman rootless
SCRAPER_PRIVATE_URL=http://someuser:somepass@scraper:3000/
LOCAL_SCRAPER_USERNAME=someuser
LOCAL_SCRAPER_PASSWORD=somepass
[!IMPORTANT] Key changes for Podman:
- Use HTTP instead of HTTPS for
SCRAPER_PRIVATE_URL- Use port 3000 instead of 443
- Change internal
exposeto3000/tcp- Update port mapping to target
3000instead of443
3. Recreate containers:
podman-compose down
podman-compose up -d --force-recreate
4. Test scraper connectivity:
# Test from within the pentagi container
podman exec -it pentagi wget -O- "http://someuser:somepass@scraper:3000/html?url=http://example.com"
If you see HTML output, the scraper is working correctly.
If you're running Podman in rootful mode (with sudo), you can use the default configuration without modifications. The scraper will work on port 443 as intended.
All Podman configurations remain fully compatible with Docker. The non-privileged port approach works identically on both container runtimes.
PentAGI allows you to configure default behavior for assistants:
| Variable | Default | Description |
|---|---|---|
ASSISTANT_USE_AGENTS | false | Controls the default value for agent usage when creating new assistants |
The ASSISTANT_USE_AGENTS setting affects the initial state of the "Use Agents" toggle when creating a new assistant in the UI:
false (default): New assistants are created with agent delegation disabled by defaulttrue: New assistants are created with agent delegation enabled by defaultNote that users can always override this setting by toggling the "Use Agents" button in the UI when creating or editing an assistant. This environment variable only controls the initial default state.
Once the stack is running and you can sign in to the web UI, the fastest way to start is through the Flows workflow.
Good first prompts usually include:
Example:
Assess https://target.example for common web application vulnerabilities. Focus on authentication, file handling, and injection issues. Stay within the provided target only and summarize confirmed findings with reproduction steps.
Only test systems you own or are explicitly authorized to assess. See EULA.md for the acceptable use requirements.
The new flow form includes a template picker, which can prefill the message box with a saved flow template. This is useful when you run similar assessments repeatedly.
examples/prompts/base_web_pentest.md if you need a practical baseline for web testingTemplates are starting points. You do not need special syntax to use PentAGI: plain natural-language instructions work well as long as the target and goal are clear.
After submitting the flow, PentAGI opens the flow page automatically.
Once the flow has enough results, use the Report menu on the flow page to:
Each flow also includes an Assistant view for interactive guidance. This is useful when the autonomous run uncovers something that needs human direction instead of a hard restart.
Each flow has its own Files tab in the flow page. Files are scoped to the parent flow: they live in {dataDir}/flow-{id}-data/ on the host and never leak into other flows.
The tab exposes three sources of files:
uploads/): files you provide from the web UI. Use the Upload files action, or drag and drop directly onto the Files tab. While the agent container is running, uploaded files are also pushed into it at /work/uploads/ so the agent can read them with normal shell tools.resources/): files attached from your saved user resources library via Attach resources from library. Attached resources are copied into the flow and pushed into the running container at /work/resources/.container/): snapshots pulled from the running agent container via Pull file or directory from container. These are read-only on the flow side and are never sent back to the container.Per-file actions in the Files tab include Download, Copy path, Save as resource (promote a flow file into your reusable resources library), and Delete. The Pull action is disabled when the container is not running, with the tooltip "Container is not running".
Uploaded files and attached resources are listed automatically in the agent's system prompts via the {{.UserFiles}} template variable, which renders a compact <task_files> XML block (with nested <uploads> and <resources> sections), so the assistant and automation agents can reference them by path without you pasting the contents into chat. Container snapshots are visible in the UI only and are not auto-injected back into the prompt.
Current limits and limitations to be aware of:
/work/uploads/ and /work/resources/; files written to other container paths are not auto-mirrored back into the flow file model. Container snapshots can originate from any container path you pull (for example /etc/...) and are cached on the flow side under container/; they are not pushed back into the container.flow-{id}-data/ directory on disk. Operators are still expected to clean up the data directory manually if they want to reclaim the space.For early testing, start with a narrow target and a single clear objective. This makes the output easier to review and helps you refine your prompts before running larger assessments.
PentAGI provides comprehensive programmatic access through both REST and GraphQL APIs, allowing you to integrate penetration testing workflows into your automation pipelines, CI/CD processes, and custom applications.
API tokens are managed through the PentAGI web interface:
Each token is associated with your user account and inherits your role's permissions.
Include the API token in the Authorization header of your HTTP requests:
# GraphQL API example
curl -X POST https://your-pentagi-instance:8443/api/v1/graphql \
-H "Authorization: Bearer YOUR_API_TOKEN" \
-H "Content-Type: application/json" \
-d '{"query": "{ flows { id title status } }"}'
# REST API example
curl https://your-pentagi-instance:8443/api/v1/flows \
-H "Authorization: Bearer YOUR_API_TOKEN"
PentAGI provides interactive documentation for exploring and testing API endpoints:
Access the GraphQL Playground at https://your-pentagi-instance:8443/api/v1/graphql/playground
{
"Authorization": "Bearer YOUR_API_TOKEN"
}
Access the REST API documentation at https://your-pentagi-instance:8443/api/v1/swagger/index.html
Bearer YOUR_API_TOKENYou can generate type-safe API clients for your preferred programming language using the schema files included with PentAGI:
The GraphQL schema is available at:
schema.graphqlsbackend/pkg/graph/schema.graphqls in the repositoryGenerate clients using tools like:
The OpenAPI specification is available at:
https://your-pentagi-instance:8443/api/v1/swagger/doc.jsonbackend/pkg/server/docs/swagger.yamlGenerate clients using:
OpenAPI Generator: https://openapi-generator.tech
openapi-generator-cli generate \
-i https://your-pentagi-instance:8443/api/v1/swagger/doc.json \
-g python \
-o ./pentagi-client
Swagger Codegen: https://github.com/swagger-api/swagger-codegen
swagger-codegen generate \
-i https://your-pentagi-instance:8443/api/v1/swagger/doc.json \
-l typescript-axios \
-o ./pentagi-client
swagger-typescript-api (TypeScript): https://github.com/acacode/swagger-typescript-api
npx swagger-typescript-api \
-p https://your-pentagi-instance:8443/api/v1/swagger/doc.json \
-o ./src/api \
-n pentagi-api.ts
mutation CreateFlow {
createFlow(
modelProvider: "openai"
input: "Test the security of https://example.com"
) {
id
title
status
createdAt
}
}
curl https://your-pentagi-instance:8443/api/v1/flows \
-H "Authorization: Bearer YOUR_API_TOKEN" \
| jq '.flows[] | {id, title, status}'
import requests
class PentAGIClient:
def __init__(self, base_url, api_token):
self.base_url = base_url
self.headers = {
"Authorization": f"Bearer {api_token}",
"Content-Type": "application/json"
}
def create_flow(self, provider, target):
query = """
mutation CreateFlow($provider: String!, $input: String!) {
createFlow(modelProvider: $provider, input: $input) {
id
title
status
}
}
"""
response = requests.post(
f"{self.base_url}/api/v1/graphql",
json={
"query": query,
"variables": {
"provider": provider,
"input": target
}
},
headers=self.headers
)
return response.json()
def get_flows(self):
response = requests.get(
f"{self.base_url}/api/v1/flows",
headers=self.headers
)
return response.json()
# Usage
client = PentAGIClient(
"https://your-pentagi-instance:8443",
"your_api_token_here"
)
# Create a new flow
flow = client.create_flow("openai", "Scan https://example.com for vulnerabilities")
print(f"Created flow: {flow}")
# List all flows
flows = client.get_flows()
print(f"Total flows: {len(flows['flows'])}")
import axios, { AxiosInstance } from 'axios';
interface Flow {
id: string;
title: string;
status: string;
createdAt: string;
}
class PentAGIClient {
private client: AxiosInstance;
constructor(baseURL: string, apiToken: string) {
this.client = axios.create({
baseURL: `${baseURL}/api/v1`,
headers: {
'Authorization': `Bearer ${apiToken}`,
'Content-Type': 'application/json',
},
});
}
async createFlow(provider: string, input: string): Promise<Flow> {
const query = `
mutation CreateFlow($provider: String!, $input: String!) {
createFlow(modelProvider: $provider, input: $input) {
id
title
status
createdAt
}
}
`;
const response = await this.client.post('/graphql', {
query,
variables: { provider, input },
});
return response.data.data.createFlow;
}
async getFlows(): Promise<Flow[]> {
const response = await this.client.get('/flows');
return response.data.flows;
}
async getFlow(flowId: string): Promise<Flow> {
const response = await this.client.get(`/flows/${flowId}`);
return response.data;
}
}
// Usage
const client = new PentAGIClient(
'https://your-pentagi-instance:8443',
'your_api_token_here'
);
// Create a new flow
const flow = await client.createFlow(
'openai',
'Perform penetration test on https://example.com'
);
console.log('Created flow:', flow);
// List all flows
const flows = await client.getFlows();
console.log(`Total flows: ${flows.length}`);
When working with API tokens:
The token list shows:
When using custom LLM providers with the LLM_SERVER_* variables, you can fine-tune the reasoning format used in requests.
[!TIP] For production-grade local deployments, consider using vLLM with Qwen3.5-27B-FP8 for optimal performance. See our comprehensive deployment guide which includes hardware requirements, configuration templates (thinking mode and non-thinking mode), and performance benchmarks showing 13K TPS prompt processing on 4× RTX 5090 GPUs.
| Variable | Default | Description |
|---|---|---|
LLM_SERVER_URL | Base URL for the custom LLM API endpoint | |
LLM_SERVER_KEY | API key for the custom LLM provider | |
LLM_SERVER_MODEL | Default model to use (can be overridden in provider config) | |
LLM_SERVER_CONFIG_PATH | Path to the YAML configuration file for agent-specific models | |
LLM_SERVER_PROVIDER | Provider name prefix for model names (e.g., openrouter, deepseek for LiteLLM proxy) | |
LLM_SERVER_LEGACY_REASONING | false | Controls reasoning format in API requests |
LLM_SERVER_PRESERVE_REASONING | false | Preserve reasoning content in multi-turn conversations (required by some providers) |
The LLM_SERVER_PROVIDER setting is particularly useful when using LiteLLM proxy, which adds a provider prefix to model names. For example, when connecting to Moonshot API through LiteLLM, models like kimi-2.5 become moonshot/kimi-2.5. By setting LLM_SERVER_PROVIDER=moonshot, you can use the same provider configuration file for both direct API access and LiteLLM proxy access without modifications.
The LLM_SERVER_LEGACY_REASONING setting affects how reasoning parameters are sent to the LLM:
false (default): Uses modern format where reasoning is sent as a structured object with max_tokens parametertrue: Uses legacy format with string-based reasoning_effort parameterThis setting is important when working with different LLM providers as they may expect different reasoning formats in their API requests. If you encounter reasoning-related errors with custom providers, try changing this setting.
The LLM_SERVER_PRESERVE_REASONING setting controls whether reasoning content is preserved in multi-turn conversations:
false (default): Reasoning content is not preserved in conversation historytrue: Reasoning content is preserved and sent in subsequent API callsThis setting is required by some LLM providers (e.g., Moonshot) that return errors like "thinking is enabled but reasoning_content is missing in assistant tool call message" when reasoning content is not included in multi-turn conversations. Enable this setting if your provider requires reasoning content to be preserved.
PentAGI supports Ollama for both local LLM inference (zero-cost, enhanced privacy) and Ollama Cloud (managed service with free tier).
| Variable | Default | Description |
|---|---|---|
OLLAMA_SERVER_URL | URL of your Ollama server or Ollama Cloud | |
OLLAMA_SERVER_API_KEY | API key for Ollama Cloud authentication | |
OLLAMA_SERVER_MODEL | Default model for inference | |
OLLAMA_SERVER_CONFIG_PATH | Path to custom agent configuration file | |
OLLAMA_SERVER_PULL_MODELS_TIMEOUT | 600 | Timeout for model downloads (seconds) |
OLLAMA_SERVER_PULL_MODELS_ENABLED | false | Auto-download models on startup |
OLLAMA_SERVER_LOAD_MODELS_ENABLED | false | Query server for available models |
Ollama Cloud provides managed inference with a generous free tier and scalable paid plans.
Free Tier Setup (Single Model)
# Free tier allows one model at a time
OLLAMA_SERVER_URL=https://ollama.com
OLLAMA_SERVER_API_KEY=your_ollama_cloud_api_key
OLLAMA_SERVER_MODEL=gpt-oss:120b # Example: OpenAI OSS 120B model
Paid Tier Setup (Multi-Model with Pre-built Configuration)
For paid tiers supporting multiple concurrent models, use the pre-built Ollama Cloud configuration:
# Using pre-built Ollama Cloud configuration (included in Docker image)
OLLAMA_SERVER_URL=https://ollama.com
OLLAMA_SERVER_API_KEY=your_ollama_cloud_api_key
OLLAMA_SERVER_CONFIG_PATH=/opt/pentagi/conf/ollama-cloud.provider.yml
The pre-built ollama-cloud.provider.yml configuration includes optimized model assignments for all agent types:
nemotron-3-super:cloud - Fast general-purpose modelqwen3-coder-next:cloud - Advanced reasoning with high effort modeqwen3-coder-next:cloud - Specialized coding modelsqwen3.5:397b-cloud - Large context for information gatheringglm-5:cloud - High-quality text refinementminimax-m2.7:cloud - Efficient advisory tasksdevstral-2:123b-cloud - Installation and setup tasksCustom Configuration (Advanced)
To create your own agent configuration, mount a custom file from your host filesystem:
# Using custom provider configuration
OLLAMA_SERVER_URL=https://ollama.com
OLLAMA_SERVER_API_KEY=your_ollama_cloud_api_key
OLLAMA_SERVER_CONFIG_PATH=/opt/pentagi/conf/ollama.provider.yml
# Mount custom configuration from host filesystem (in .env or docker-compose override)
PENTAGI_OLLAMA_SERVER_CONFIG_PATH=/path/on/host/my-ollama-config.yml
The PENTAGI_OLLAMA_SERVER_CONFIG_PATH environment variable maps your host configuration file to /opt/pentagi/conf/ollama.provider.yml inside the container.
Example custom configuration (my-ollama-config.yml):
primary_agent:
model: "qwen3-coder-next:cloud"
temperature: 1.0
top_p: 0.9
max_tokens: 32768
reasoning:
effort: high
coder:
model: "qwen3-coder:32b"
temperature: 1.0
max_tokens: 20480
For self-hosted Ollama instances:
# Basic local Ollama setup
OLLAMA_SERVER_URL=http://localhost:11434
OLLAMA_SERVER_MODEL=llama3.1:8b-instruct-q8_0
# Production setup with auto-pull and model discovery
OLLAMA_SERVER_URL=http://ollama-server:11434
OLLAMA_SERVER_PULL_MODELS_ENABLED=true
OLLAMA_SERVER_PULL_MODELS_TIMEOUT=900
OLLAMA_SERVER_LOAD_MODELS_ENABLED=true
# Using pre-built configurations from Docker image
OLLAMA_SERVER_CONFIG_PATH=/opt/pentagi/conf/ollama-llama318b.provider.yml
# or
OLLAMA_SERVER_CONFIG_PATH=/opt/pentagi/conf/ollama-qwen332b-fp16-tc.provider.yml
# or
OLLAMA_SERVER_CONFIG_PATH=/opt/pentagi/conf/ollama-qwq32b-fp16-tc.provider.yml
Performance Considerations:
OLLAMA_SERVER_LOAD_MODELS_ENABLED=true): Adds 1-2s startup latency querying Ollama APIOLLAMA_SERVER_PULL_MODELS_ENABLED=true): First startup may take several minutes downloading modelsOLLAMA_SERVER_PULL_MODELS_TIMEOUT=900): 15 minutes in secondsPentAGI requires models with larger context windows than the default Ollama configurations. You need to create custom models with increased num_ctx parameter through Modelfiles. While typical agent workflows consume around 64K tokens, PentAGI uses 110K context size for safety margin and handling complex penetration testing scenarios.
Important: The num_ctx parameter can only be set during model creation via Modelfile - it cannot be changed after model creation or overridden at runtime.
Create a Modelfile named Modelfile_qwen3_32b_fp16_tc:
FROM qwen3:32b-fp16
PARAMETER num_ctx 110000
PARAMETER temperature 0.3
PARAMETER top_p 0.8
PARAMETER min_p 0.0
PARAMETER top_k 20
PARAMETER repeat_penalty 1.1
Build the custom model:
ollama create qwen3:32b-fp16-tc -f Modelfile_qwen3_32b_fp16_tc
Create a Modelfile named Modelfile_qwq_32b_fp16_tc:
FROM qwq:32b-fp16
PARAMETER num_ctx 110000
PARAMETER temperature 0.2
PARAMETER top_p 0.7
PARAMETER min_p 0.0
PARAMETER top_k 40
PARAMETER repeat_penalty 1.2
Build the custom model:
ollama create qwq:32b-fp16-tc -f Modelfile_qwq_32b_fp16_tc
Note: The QwQ 32B FP16 model requires approximately 71.3 GB VRAM for inference. Ensure your system has sufficient GPU memory before attempting to use this model.
These custom models are referenced in the pre-built provider configuration files (ollama-qwen332b-fp16-tc.provider.yml and ollama-qwq32b-fp16-tc.provider.yml) that are included in the Docker image at /opt/pentagi/conf/.
PentAGI integrates with OpenAI's comprehensive model lineup, featuring advanced reasoning capabilities with extended chain-of-thought, agentic models with enhanced tool integration, and specialized code models for security engineering.
| Variable | Default | Description |
|---|---|---|
OPEN_AI_KEY | API key for OpenAI services | |
OPEN_AI_SERVER_URL | https://api.openai.com/v1 | OpenAI API endpoint |
# Basic OpenAI setup
OPEN_AI_KEY=your_openai_api_key
OPEN_AI_SERVER_URL=https://api.openai.com/v1
# Using with proxy for enhanced security
OPEN_AI_KEY=your_openai_api_key
PROXY_URL=http://your-proxy:8080
PentAGI supports 31 OpenAI models with tool calling, streaming, reasoning modes, and prompt caching. Models marked with * are used in default configuration.
GPT-5.2 Series - Latest Flagship Agentic (December 2025)
| Model ID | Thinking | Price (Input/Output/Cache) | Use Case |
|---|---|---|---|
gpt-5.2* | ✅ | $1.75/$14.00/$0.18 | Latest flagship with enhanced reasoning and tool integration, autonomous security research |
gpt-5.2-pro | ✅ | $21.00/$168.00/$0.00 | Premium version with superior agentic coding, mission-critical security research, zero-day discovery |
gpt-5.2-codex | ✅ | $1.75/$14.00/$0.18 | Most advanced code-specialized, context compaction, strong cybersecurity capabilities |
GPT-5/5.1 Series - Advanced Agentic Models
| Model ID | Thinking | Price (Input/Output/Cache) | Use Case |
|---|---|---|---|
gpt-5 | ✅ | $1.25/$10.00/$0.13 | Premier agentic with advanced reasoning, autonomous security research, exploit chain development |
gpt-5.1 | ✅ | $1.25/$10.00/$0.13 | Enhanced agentic with adaptive reasoning, balanced penetration testing with strong tool coordination |
gpt-5-pro | ✅ | $15.00/$120.00/$0.00 | Premium version with major reasoning improvements, reduced hallucinations, critical security operations |
gpt-5-mini | ✅ | $0.25/$2.00/$0.03 | Efficient balancing speed and intelligence, automated vulnerability analysis, exploit generation |
gpt-5-nano | ✅ | $0.05/$0.40/$0.01 | Fastest for high-throughput scanning, reconnaissance, bulk vulnerability detection |
GPT-5/5.1 Codex Series - Code-Specialized
| Model ID | Thinking | Price (Input/Output/Cache) | Use Case |
|---|---|---|---|
gpt-5.1-codex-max | ✅ | $1.25/$10.00/$0.13 | Enhanced reasoning for sophisticated coding, proven CVE findings, systematic exploit development |
gpt-5.1-codex | ✅ | $1.25/$10.00/$0.13 | Standard code-optimized with strong reasoning, exploit generation, vulnerability analysis |
gpt-5-codex | ✅ | $1.25/$10.00/$0.13 | Foundational code-specialized, vulnerability scanning, basic exploit generation |
gpt-5.1-codex-mini | ✅ | $0.25/$2.00/$0.03 | Compact high-performance, 4x higher capacity, rapid vulnerability detection |
codex-mini-latest | ✅ | $1.50/$6.00/$0.38 | Latest compact code model, automated code review, basic vulnerability analysis |
GPT-4.1 Series - Enhanced Intelligence
| Model ID | Thinking | Price (Input/Output/Cache) | Use Case |
|---|---|---|---|
gpt-4.1 | ❌ | $2.00/$8.00/$0.50 | Enhanced flagship with superior function calling, complex threat analysis, sophisticated exploit development |
gpt-4.1-mini* | ❌ | $0.40/$1.60/$0.10 | Balanced performance with improved efficiency, routine security assessments, automated code analysis |
gpt-4.1-nano | ❌ | $0.10/$0.40/$0.03 | Ultra-fast lightweight, bulk security scanning, rapid reconnaissance, continuous monitoring |
GPT-4o Series - Multimodal Flagship
| Model ID | Thinking | Price (Input/Output/Cache) | Use Case |
|---|---|---|---|
gpt-4o | ❌ | $2.50/$10.00/$1.25 | Multimodal flagship with vision, image analysis, web UI assessment, multi-tool orchestration |
gpt-4o-mini | ❌ | $0.15/$0.60/$0.08 | Compact multimodal with strong function calling, high-frequency scanning, cost-effective bulk operations |
o-Series - Advanced Reasoning Models
| Model ID | Thinking | Price (Input/Output/Cache) | Use Case |
|---|---|---|---|
o4-mini* | ✅ | $1.10/$4.40/$0.28 | Next-gen reasoning with enhanced speed, methodical security assessments, systematic exploit development |
o3* | ✅ | $2.00/$8.00/$0.50 | Advanced reasoning powerhouse, multi-stage attack chains, deep vulnerability analysis |
o3-mini | ✅ | $1.10/$4.40/$0.55 | Compact reasoning with extended thinking, step-by-step attack planning, logical vulnerability chaining |
o1 | ✅ | $15.00/$60.00/$7.50 | Premier reasoning with maximum depth, advanced penetration testing, novel exploit research |
o3-pro | ✅ | $20.00/$80.00/$0.00 | Most advanced reasoning, 80% cheaper than o1-pro, zero-day research, critical security investigations |
o1-pro | ✅ | $150.00/$600.00/$0.00 | Previous-gen premium reasoning, exhaustive security analysis, mission-critical challenges |
Prices: Per 1M tokens. Reasoning models include thinking tokens in output pricing.
[!WARNING] GPT-5 Models - Trusted Access Required*
All GPT-5 series models (
gpt-5,gpt-5.1,gpt-5.2,gpt-5-pro,gpt-5.2-pro, and all Codex variants) work unstably with PentAGI and may trigger OpenAI's cybersecurity safety mechanisms without verified access.To use GPT-5 models reliably:*
- Individual users: Verify your identity at chatgpt.com/cyber
- Enterprise teams: Request trusted access through your OpenAI representative
- Security researchers: Apply for the Cybersecurity Grant Program (includes $10M in API credits)
Recommended alternatives without verification:
- Use
o-seriesmodels (o3, o4-mini, o1) for reasoning tasks- Use
gpt-4.1series for general intelligence and function calling- All o-series and gpt-4.x models work reliably without special access
Reasoning Effort Levels:
Key Features:
PentAGI integrates with Anthropic's Claude models, featuring advanced extended thinking capabilities, exceptional safety mechanisms, and sophisticated understanding of complex security contexts with prompt caching.
| Variable | Default | Description |
|---|---|---|
ANTHROPIC_API_KEY | API key for Anthropic services | |
ANTHROPIC_SERVER_URL | https://api.anthropic.com/v1 | Anthropic API endpoint |
# Basic Anthropic setup
ANTHROPIC_API_KEY=your_anthropic_api_key
ANTHROPIC_SERVER_URL=https://api.anthropic.com/v1
# Using with proxy for secure environments
ANTHROPIC_API_KEY=your_anthropic_api_key
PROXY_URL=http://your-proxy:8080
[!NOTE] Google Vertex AI for Claude models
PentAGI does not currently expose a dedicated Google Vertex AI configuration path for Anthropic Claude in
.env. There is no separate Vertex AI API key field at this time, and the existing Anthropic variables (ANTHROPIC_API_KEY,ANTHROPIC_SERVER_URL) target the direct Anthropic API. Supported routes for Claude are:
- Direct Anthropic API:
ANTHROPIC_API_KEYandANTHROPIC_SERVER_URL(see above).- AWS Bedrock:
BEDROCK_*variables (see AWS Bedrock Provider Configuration).If you need to use Vertex AI today, the safest supported workaround is to expose Vertex AI through an OpenAI-compatible proxy or gateway that translates Vertex AI calls into the Chat Completions format while preserving the chat and tool-call behavior PentAGI relies on, then point the Custom LLM provider at that gateway via
LLM_SERVER_URL,LLM_SERVER_KEY, andLLM_SERVER_MODEL. This path is only as reliable as the gateway you choose.
PentAGI supports 10 Claude models with tool calling, streaming, extended thinking, adaptive thinking, and prompt caching. Models marked with * are used in default configuration.
Claude 4 Series - Latest Models (2025-2026)
| Model ID | Thinking | Release Date | Price (Input/Output/Cache R/W) | Use Case |
|---|---|---|---|---|
claude-opus-4-6* | ✅ | May 2025 | $5.00/$25.00/$0.50/$6.25 | Most intelligent model for autonomous agents and coding. Extended + adaptive thinking for complex exploit development, multi-stage attack simulation |
claude-sonnet-4-6* | ✅ | Aug 2025 | $3.00/$15.00/$0.30/$3.75 | Best speed/intelligence balance with adaptive thinking. Multi-phase security assessments, intelligent vulnerability analysis, real-time threat hunting |
claude-haiku-4-5* | ✅ | Oct 2025 | $1.00/$5.00/$0.10/$1.25 | Fastest model with near-frontier intelligence. High-frequency scanning, real-time monitoring, bulk automated testing |
Legacy Models - Still Supported
| Model ID | Thinking | Release Date | Price (Input/Output/Cache R/W) | Use Case |
|---|---|---|---|---|
claude-sonnet-4-5 | ✅ | Sep 2025 | $3.00/$15.00/$0.30/$3.75 | State-of-the-art reasoning (superseded by 4-6). Sophisticated penetration testing, advanced threat analysis |
claude-opus-4-5 | ✅ | Nov 2025 | $5.00/$25.00/$0.50/$6.25 | Ultimate reasoning (superseded by opus-4-6). Critical security research, zero-day discovery, red team operations |
claude-opus-4-1 | ✅ | Aug 2025 | $15.00/$75.00/$1.50/$18.75 | Advanced reasoning (superseded). Complex penetration testing, sophisticated threat modeling |
claude-sonnet-4-0 | ✅ | May 2025 | $3.00/$15.00/$0.30/$3.75 | High-performance reasoning (superseded). Complex threat modeling, multi-tool coordination |
claude-opus-4-0 | ✅ | May 2025 | $15.00/$75.00/$1.50/$18.75 | First generation Opus (superseded). Multi-step exploit development, autonomous pentesting workflows |
Deprecated Models - Migrate to Current Models
| Model ID | Thinking | Release Date | Price (Input/Output/Cache R/W) | Notes |
|---|---|---|---|---|
claude-3-haiku-20240307 | ❌ | Mar 2024 | $0.25/$1.25/$0.03/$0.30 | Will be retired April 19, 2026. Migrate to claude-haiku-4-5 |
Prices: Per 1M tokens. Cache pricing includes both Read and Write costs.
Extended Thinking Configuration:
Key Features:
PentAGI integrates with Google's Gemini models through the Google AI API, offering state-of-the-art multimodal reasoning capabilities with extended thinking and context caching.
| Variable | Default | Description |
|---|---|---|
GEMINI_API_KEY | API key for Google AI services | |
GEMINI_SERVER_URL | https://generativelanguage.googleapis.com | Google AI API endpoint |
# Basic Gemini setup
GEMINI_API_KEY=your_gemini_api_key
GEMINI_SERVER_URL=https://generativelanguage.googleapis.com
# Using with proxy
GEMINI_API_KEY=your_gemini_api_key
PROXY_URL=http://your-proxy:8080
PentAGI supports 9 Gemini models with tool calling, streaming, thinking modes, and context caching. Models marked with * are used in default configuration.
Gemini 3.5 Series - Latest Stable Flash (May 2026)
| Model ID | Thinking | Context | Price (Input/Output/Cache) | Use Case |
|---|---|---|---|---|
gemini-3.5-flash* | ✅ | 1M | $1.50/$9.00/$0.15 | Most intelligent Flash model with sustained frontier performance on agentic and coding tasks, superior search and grounding |
Gemini 3.1 Series - Stable Flash-Lite + Pro Preview (Feb-May 2026)
| Model ID | Thinking | Context | Price (Input/Output/Cache) | Use Case |
|---|---|---|---|---|
gemini-3.1-pro-preview* | ✅ | 1M | $2.00/$12.00/$0.20 | Latest flagship with refined thinking, improved token efficiency, optimized for software engineering and agentic workflows |
gemini-3.1-pro-preview-customtools | ✅ | 1M | $2.00/$12.00/$0.20 | Custom tools endpoint optimized for bash and custom tools (view_file, search_code) prioritization |
gemini-3.1-flash-lite* | ✅ | 1M | $0.25/$1.50/$0.025 | Most cost-efficient stable multimodal model, frontier-class performance for high-volume agentic tasks and low-latency applications |
Gemini 2.5 Series - Advanced Thinking Models (active until October 16, 2026)
| Model ID | Thinking | Context | Price (Input/Output/Cache) | Use Case |
|---|---|---|---|---|
gemini-2.5-pro | ✅ | 1M | $1.25/$10.00/$0.125 | State-of-the-art for complex coding and reasoning, sophisticated threat modeling |
gemini-2.5-flash | ✅ | 1M | $0.30/$2.50/$0.03 | First hybrid reasoning model with thinking budgets, best price-performance for large-scale assessments |
gemini-2.5-flash-lite | ✅ | 1M | $0.10/$0.40/$0.01 | Smallest and most cost-effective for at-scale usage, high-throughput scanning |
Gemma 4 Open-Source Models (Apache 2.0, Free Tier)
| Model ID | Thinking | Context | Price (Input/Output/Cache) | Use Case |
|---|---|---|---|---|
gemma-4-31b-it | ✅ | 256K | Free/Free/Free | Largest open-source Gemma 4 dense model (~31B params), multimodal text+image, 140+ languages, on-premises security operations |
gemma-4-26b-a4b-it | ✅ | 256K | Free/Free/Free | MoE architecture (~26B total / ~3.8B active params), highly efficient inference on consumer GPUs for on-premises high-throughput scanning |
Prices: Per 1M tokens (Standard Paid tier). Context window is input token limit.
[!NOTE] Gemini 2.5 Series Shutdown
gemini-2.5-pro,gemini-2.5-flash, andgemini-2.5-flash-litewill be shut down on October 16, 2026. Recommended migrations:
gemini-2.5-pro→gemini-3.1-pro-preview(same $2.00 input pricing tier)gemini-2.5-flash→gemini-3.5-flash(improved frontier capabilities)gemini-2.5-flash-lite→gemini-3.1-flash-lite(same $0.25 input pricing)
Default Model Assignments (config.yml):
gemini-3.1-pro-preview - primary_agent, assistant, generator, refiner, adviser, coder, pentestergemini-3.5-flash - reflector, searcher, enricher, installergemini-3.1-flash-lite - simple, simple_jsonKey Features:
gemini-3.1-pro-preview-customtools route for tool-heavy agentic workflows that prefer registered tools over bashReasoning Effort Levels:
PentAGI integrates with Amazon Bedrock, offering access to 20+ foundation models from leading AI companies including Anthropic, Amazon, Cohere, DeepSeek, OpenAI, Qwen, Mistral, and Moonshot.
| Variable | Default | Description |
|---|---|---|
BEDROCK_REGION | us-east-1 | AWS region for Bedrock service |
BEDROCK_DEFAULT_AUTH | false | Use AWS SDK default credential chain (environment, EC2 role, ~/.aws/credentials) - highest priority |
BEDROCK_BEARER_TOKEN | Bearer token authentication - priority over static credentials | |
BEDROCK_ACCESS_KEY_ID | AWS access key ID for static credentials | |
BEDROCK_SECRET_ACCESS_KEY | AWS secret access key for static credentials | |
BEDROCK_SESSION_TOKEN | AWS session token for temporary credentials (optional, used with static credentials) | |
BEDROCK_SERVER_URL | Custom Bedrock endpoint (VPC endpoints, local testing) |
Authentication Priority: BEDROCK_DEFAULT_AUTH → BEDROCK_BEARER_TOKEN → BEDROCK_ACCESS_KEY_ID+BEDROCK_SECRET_ACCESS_KEY
# Recommended: Default AWS SDK authentication (EC2/ECS/Lambda roles)
BEDROCK_REGION=us-east-1
BEDROCK_DEFAULT_AUTH=true
# Bearer token authentication (AWS STS, custom auth)
BEDROCK_REGION=us-east-1
BEDROCK_BEARER_TOKEN=your_bearer_token
# Static credentials (development, testing)
BEDROCK_REGION=us-east-1
BEDROCK_ACCESS_KEY_ID=your_aws_access_key
BEDROCK_SECRET_ACCESS_KEY=your_aws_secret_key
# With proxy and custom endpoint
BEDROCK_REGION=us-east-1
BEDROCK_DEFAULT_AUTH=true
BEDROCK_SERVER_URL=https://bedrock-runtime.us-east-1.vpce-xxx.amazonaws.com
PROXY_URL=http://your-proxy:8080
PentAGI supports 21 AWS Bedrock models with tool calling, streaming, and multimodal capabilities. Models marked with * are used in default configuration.
| Model ID | Provider | Thinking | Multimodal | Price (Input/Output) | Use Case |
|---|---|---|---|---|---|
us.amazon.nova-2-lite-v1:0 | Amazon Nova | ❌ | ✅ | $0.33/$2.75 | Adaptive reasoning, efficient thinking |
us.amazon.nova-premier-v1:0 | Amazon Nova | ❌ | ✅ | $2.50/$12.50 | Complex reasoning, advanced analysis |
us.amazon.nova-pro-v1:0 | Amazon Nova | ❌ | ✅ | $0.80/$3.20 | Balanced accuracy, speed, cost |
us.amazon.nova-lite-v1:0 | Amazon Nova | ❌ | ✅ | $0.06/$0.24 | Fast processing, high-volume operations |
us.amazon.nova-micro-v1:0 | Amazon Nova | ❌ | ❌ | $0.035/$0.14 | Ultra-low latency, real-time monitoring |
us.anthropic.claude-opus-4-6-v1* | Anthropic | ✅ | ✅ | $5.00/$25.00 | World-class coding, enterprise agents |
us.anthropic.claude-sonnet-4-6 | Anthropic | ✅ | ✅ | $3.00/$15.00 | Frontier intelligence, enterprise scale |
us.anthropic.claude-opus-4-5-20251101-v1:0 | Anthropic | ✅ | ✅ | $5.00/$25.00 | Multi-day software development |
us.anthropic.claude-haiku-4-5-20251001-v1:0* | Anthropic | ✅ | ✅ | $1.00/$5.00 | Near-frontier performance, high speed |
us.anthropic.claude-sonnet-4-5-20250929-v1:0* | Anthropic | ✅ | ✅ | $3.00/$15.00 | Real-world agents, coding excellence |
us.anthropic.claude-sonnet-4-20250514-v1:0 | Anthropic | ✅ | ✅ | $3.00/$15.00 | Balanced performance, production-ready |
us.anthropic.claude-3-5-haiku-20241022-v1:0 | Anthropic | ❌ | ❌ | $0.80/$4.00 | Fastest model, cost-effective scanning |
Prices: Per 1M tokens. Models with thinking/reasoning support additional compute costs during reasoning phase.
Some AWS Bedrock models were tested but are not supported due to technical limitations:
| Model Family | Reason for Incompatibility |
|---|---|
| GLM (Z.AI) | Tool calling format incompatible with Converse API (expects string instead of JSON) |
| AI21 Jamba | Severe rate limits (1-2 req/min) prevent reliable testing and production use |
| Meta Llama 3.3/3.1 | Unstable tool call result processing, causes unexpected failures in multi-turn workflows |
| Mistral Magistral | Tool calling not supported by the model |
| Moonshot K2-Thinking | Unstable streaming behavior with tool calls, unreliable in production |
| Qwen3-VL | Unstable streaming with tool calling, multimodal + tools combination fails intermittently |
[!IMPORTANT] Rate Limits & Quota Management
Default AWS Bedrock quotas for Claude models are extremely restrictive (2-20 requests/minute for new accounts). For production penetration testing:
- Request quota increases through AWS Service Quotas console for models you plan to use
- Use Amazon Nova models - higher default quotas and excellent performance
- Enable provisioned throughput for consistent high-volume testing
- Monitor usage - AWS throttles aggressively at quota limits
Without quota increases, expect frequent delays and workflow interruptions.
[!WARNING] Converse API Requirements
PentAGI uses Amazon Bedrock Converse API for unified model access. All supported models require:
- ✅ Converse/ConverseStream API support
- ✅ Tool use (function calling) for penetration testing workflows
- ✅ Streaming tool use for real-time feedback
Verify model capabilities at: AWS Bedrock Model Features
Key Features:
PentAGI integrates with DeepSeek, providing access to advanced AI models with strong reasoning, coding capabilities, and context caching at competitive prices.
| Variable | Default Value | Description |
|---|---|---|
DEEPSEEK_API_KEY | DeepSeek API key for authentication | |
DEEPSEEK_SERVER_URL | https://api.deepseek.com | DeepSeek API endpoint URL |
DEEPSEEK_PROVIDER | Provider prefix for LiteLLM integration (optional) |
# Direct API usage
DEEPSEEK_API_KEY=your_deepseek_api_key
DEEPSEEK_SERVER_URL=https://api.deepseek.com
# With LiteLLM proxy
DEEPSEEK_API_KEY=your_litellm_key
DEEPSEEK_SERVER_URL=http://litellm-proxy:4000
DEEPSEEK_PROVIDER=deepseek # Adds prefix to model names (deepseek/deepseek-v4-flash) for LiteLLM
PentAGI supports 2 DeepSeek V4 models with tool calling, streaming, hybrid thinking/non-thinking modes, and context caching. Both models support thinking mode by default and can be switched to non-thinking mode via extra_body. Models marked with * are used in default configuration.
| Model ID | Thinking | Max Output | Context | Price (Input/Output/Cache) | Use Case |
|---|---|---|---|---|---|
deepseek-v4-flash* | ✅ hybrid | 384K | 1M | $0.14/$0.28/$0.0028 | Utility agents, general dialogue, fast tool calling |
deepseek-v4-pro* | ✅ hybrid | 384K | 1M | $1.74/$3.48/$0.0145 | Advanced reasoning, complex logic, security analysis |
Prices: Per 1M tokens. Cache pricing applies to prompt tokens served from cache (input cache hit, reduced to 1/10 of launch price since 2026-04-26). Both models support hybrid thinking — thinking mode is enabled by default; pass extra_body.thinking.type: disabled to switch to non-thinking mode for faster/cheaper responses.
Pricing Note (deepseek-v4-pro): The 75% promotional discount on
deepseek-v4-proofficially ended on 2026-05-31 15:59 UTC. The prices above reflect the standard post-promotional pricing. If you have legacy configurations using the discounted prices ($0.435/$0.87/$0.003625), update them to the current rates for accurate cost tracking.
The legacy model names
deepseek-chatanddeepseek-reasonerare scheduled for deprecation by DeepSeek on 2026-07-24. Existing user configurations referencing the legacy names continue to work until then; the defaults above use the current V4 names.deepseek-chatmaps todeepseek-v4-flashnon-thinking mode;deepseek-reasonermaps todeepseek-v4-flashthinking mode.
Default Agent Configuration:
Strategy: prefer deepseek-v4-flash (12x cheaper input, 12x cheaper output) as the workhorse for utility/lightweight agents; reserve deepseek-v4-pro for complex multi-step reasoning. The installer agent runs on Flash with thinking enabled because environment setup tasks (shell commands, config edits) rarely require pro-level reasoning. Run A/B tests on your own workloads before promoting more agents to Pro.
| Agent Role | Default Model | Thinking | Reasoning Effort | Max Output | Temperature | Top P |
|---|---|---|---|---|---|---|
| Generator / Refiner | deepseek-v4-pro | Enabled | High | 32768 | (auto) | (auto) |
| Coder | deepseek-v4-pro | Enabled | High | 20480 | (auto) | (auto) |
| Primary Agent / Assistant / Pentester | deepseek-v4-pro | Enabled | High | 16384 | (auto) | (auto) |
| Adviser (mentor/planner) | deepseek-v4-pro | Enabled | High | 8192 | (auto) | (auto) |
| Installer | deepseek-v4-flash | Enabled | High | 12288 | (auto) | (auto) |
| Reflector / Searcher / Enricher | deepseek-v4-flash | Disabled | — | 4096 | 0.5 | 0.9 |
| Simple / Simple JSON | deepseek-v4-flash | Disabled | — | 2048 | 0.3 | 0.9 |
Note: When thinking mode is enabled, DeepSeek silently ignores
temperature,top_p,presence_penalty, andfrequency_penalty. The langchaingo client automatically nullifiestemperature/top_pwhenreasoning_effortis set, so they appear as "(auto)" in the table above. All thinking-enabled agents also explicitly passextra_body.thinking.type: enabledas defensive coding against future provider default changes.
Key Features:
extra_body.thinking.typeConcurrency Limits: deepseek-v4-flash: 2500 concurrent requests; deepseek-v4-pro: 500 concurrent requests.
LiteLLM Integration: Set DEEPSEEK_PROVIDER=deepseek to enable model name prefixing when using default PentAGI configurations with LiteLLM proxy. Leave empty for direct API usage.
PentAGI integrates with GLM from Zhipu AI (Z.AI), providing advanced language models with MoE architecture, strong reasoning, and agentic capabilities developed by Tsinghua University.
| Variable | Default Value | Description |
|---|---|---|
GLM_API_KEY | GLM API key for authentication | |
GLM_SERVER_URL | https://api.z.ai/api/paas/v4 | GLM API endpoint URL (international) |
GLM_PROVIDER | Provider prefix for LiteLLM integration (optional) |
# Direct API usage (international endpoint)
GLM_API_KEY=your_glm_api_key
GLM_SERVER_URL=https://api.z.ai/api/paas/v4
# Alternative endpoints
GLM_SERVER_URL=https://open.bigmodel.cn/api/paas/v4 # China
GLM_SERVER_URL=https://api.z.ai/api/coding/paas/v4 # Coding-specific
# With LiteLLM proxy
GLM_API_KEY=your_litellm_key
GLM_SERVER_URL=http://litellm-proxy:4000
GLM_PROVIDER=zai # Adds prefix to model names (zai/glm-4) for LiteLLM
PentAGI supports 13 GLM models with tool calling, streaming, hybrid thinking modes, and prompt caching. Models marked with * are used in default configuration. Thinking is controlled via extra_body.thinking.type ("enabled"/"disabled"); unlike Kimi, GLM is permissive about temperature in either mode.
GLM-5.x Series - Latest Generation (200K context, 128K max output)
| Model ID | Thinking | Context | Max Output | Price (Input/Output/Cache) | Use Case |
|---|---|---|---|---|---|
glm-5.1* | ✅ Hybrid | 200K | 128K | $1.40/$4.40/$0.26 | Newest flagship: 8h sustained autonomous execution, Claude Opus 4.6-aligned (generator/refiner/adviser/coder/pentester default) |
glm-5 | ✅ Hybrid | 200K | 128K | $1.00/$3.20/$0.20 | Foundation for Agentic Engineering, MoE 744B/40B active, Claude Opus 4.5-level coding |
glm-5-turbo* | ✅ Hybrid | 200K | 128K | $1.20/$4.00/$0.24 | OpenClaw-native: optimized for tool invocation, persistent tasks, long-chain execution (primary_agent/assistant default) |
GLM-4.7 Series - Premium with Interleaved Thinking
| Model ID | Thinking | Context | Max Output | Price (Input/Output/Cache) | Use Case |
|---|---|---|---|---|---|
glm-4.7 | ✅ Hybrid | 200K | 128K | $0.60/$2.20/$0.11 | Enhanced programming, stable multi-step reasoning |
glm-4.7-flashx | ✅ Hybrid | 200K | 128K | $0.07/$0.40/$0.01 | Ultra-cheap with priority GPU, but lower RPM limits (avoid for high-frequency use) |
glm-4.7-flash | ✅ Hybrid | 200K | 128K | Free/Free/Free | Free ~30B SOTA model, 1 concurrent request |
GLM-4.6 Series - Balanced with Auto-Thinking
| Model ID | Thinking | Context | Max Output | Price (Input/Output/Cache) | Use Case |
|---|---|---|---|---|---|
glm-4.6 | ✅ Auto | 200K | 128K | $0.60/$2.20/$0.11 | Balanced, streaming tool calls, token-efficient |
GLM-4.5 Series - Unified Reasoning/Coding/Agents
| Model ID | Thinking | Context | Max Output | Price (Input/Output/Cache) | Use Case |
|---|---|---|---|---|---|
glm-4.5 | ✅ Auto | 128K | 96K | $0.60/$2.20/$0.11 | Unified, MoE 355B/32B active |
glm-4.5-x | ✅ Auto | 128K | 96K | $2.20/$8.90/$0.45 | Ultra-fast premium, lowest latency |
glm-4.5-air* | ✅ Auto | 128K | 96K | $0.20/$1.10/$0.03 | Cost-effective MoE 106B/12B (simple/simple_json/reflector/searcher/enricher/installer default) |
glm-4.5-airx | ✅ Auto | 128K | 96K | $1.10/$4.50/$0.22 | Accelerated Air with priority GPU |
glm-4.5-flash | ✅ Auto | 128K | 96K | Free/Free/Free | Free with reasoning/coding/agents support |
GLM-4 Legacy - Dense Architecture
| Model ID | Thinking | Context | Max Output | Price (Input/Output) | Use Case |
|---|---|---|---|---|---|
glm-4-32b-0414-128k | ❌ | 128K | 16K | $0.10/$0.10 | Ultra-budget dense 32B, parsing without reasoning |
Prices: Per 1M tokens. Cache pricing is for prompt cache hit; cache storage is currently free per Z.AI promotion. GLM-4-32B has no cache support.
Default Agent Configuration:
Strategy: glm-5.1 (newest flagship, $1.40 input) for critical reasoning, glm-5-turbo (OpenClaw-native, agent-optimized) for orchestration, glm-4.5-air (cheap MoE with hybrid thinking and reliable RPM) for all utility/installer agents. glm-4.7-flashx is avoided as default due to lower RPM limits causing frequent 429 errors at high frequency.
| Agent Role | Default Model | Thinking | Temperature | Top P | Max Output |
|---|---|---|---|---|---|
| Generator / Refiner | glm-5.1 | Enabled | 1.0 | 0.95 | 32768 |
| Coder | glm-5.1 | Enabled | 1.0 | 0.95 | 20480 |
| Adviser / Pentester | glm-5.1 | Enabled | 1.0 | 0.95 | 16384 |
| Primary Agent / Assistant | glm-5-turbo | Enabled | 1.0 | 0.95 | 16384 |
| Installer | glm-4.5-air | Enabled | 1.0 | 0.95 | 16384 |
| Simple / Reflector | glm-4.5-air | Disabled | 0.6 | 0.9 | 8192 |
| Searcher / Enricher / Simple JSON | glm-4.5-air | Disabled | 0.6 | 0.9 | 4096 |
Note on temperature: GLM accepts both
1.0and0.6in either thinking/non-thinking mode (per Z.AI docs). langchaingo'sIsReasoningModelmatchesglm-4.5*/glm-4.6*/glm-4.7*prefixes and force-overrides temperature to 1.0 increateChatRequest— this is harmless for GLM (unlike Kimi) but means temperature values for those models in YAML are advisory.glm-5/glm-5.1/glm-5-turboare not matched, so explicit values pass through unchanged.
Thinking Modes:
extra_body.thinking.typeextra_body.thinking.clear_thinking: false so that reasoning_content from previous assistant turns is retained across the conversation. This is required on the standard API endpoint (/api/paas/v4) — on the Coding Plan endpoint it would be enabled by default. Improves reasoning continuity and cache hit rates in multi-turn tool call chains.extra_body.tool_choice: auto defensivelyKey Features:
LiteLLM Integration: Set GLM_PROVIDER=zai to enable model name prefixing when using default PentAGI configurations with LiteLLM proxy. Leave empty for direct API usage.
PentAGI integrates with Kimi from Moonshot AI, providing ultra-long context models with multimodal capabilities perfect for analyzing extensive codebases and documentation.
| Variable | Default Value | Description |
|---|---|---|
KIMI_API_KEY | Kimi API key for authentication | |
KIMI_SERVER_URL | https://api.moonshot.ai/v1 | Kimi API endpoint URL (international) |
KIMI_PROVIDER | Provider prefix for LiteLLM integration (optional) |
# Direct API usage (international endpoint)
KIMI_API_KEY=your_kimi_api_key
KIMI_SERVER_URL=https://api.moonshot.ai/v1
# Alternative endpoint
KIMI_SERVER_URL=https://api.moonshot.cn/v1 # China
# With LiteLLM proxy
KIMI_API_KEY=your_litellm_key
KIMI_SERVER_URL=http://litellm-proxy:4000
KIMI_PROVIDER=moonshot # Adds prefix to model names (moonshot/kimi-k2.5) for LiteLLM
PentAGI supports 8 Kimi/Moonshot models with tool calling, streaming, hybrid thinking modes, and multimodal capabilities (text/image/video for K2.x). All kimi-k2-* legacy models (turbo-preview, 0905-preview, 0711-preview, thinking, thinking-turbo) were deprecated by Moonshot on 2026-05-25 and are NOT included. Models marked with * are used in default configuration.
Kimi K2.x Series - Multimodal Flagship
| Model ID | Thinking | Multimodal | Context | Price (Input Miss / Output / Cache Hit) | Use Case |
|---|---|---|---|---|---|
kimi-k2.6* | ✅ hybrid | ✅ | 256K | $0.95 / $4.00 / $0.16 | Latest flagship: native multimodal, stronger code, improved instruction compliance (generator/refiner/adviser/coder/pentester default) |
kimi-k2.5* | ✅ hybrid | ✅ | 256K | $0.60 / $3.00 / $0.10 | Previous-gen: 36% cheaper input, same architecture (primary/assistant/installer/utility default) |
Moonshot V1 Series - Generation Models (Flexible Parameters)
| Model ID | Thinking | Multimodal | Context | Price (Input / Output) | Use Case |
|---|---|---|---|---|---|
moonshot-v1-8k | ❌ | ❌ | 8K | $0.20 / $2.00 | Short text generation, ultra-cheap |
moonshot-v1-32k | ❌ | ❌ | 32K | $1.00 / $3.00 | Long text generation |
moonshot-v1-128k | ❌ | ❌ | 128K | $2.00 / $5.00 | Very long context |
Moonshot V1 Vision Series - Image Understanding
| Model ID | Thinking | Multimodal | Context | Price (Input / Output) | Use Case |
|---|---|---|---|---|---|
moonshot-v1-8k-vision-preview | ❌ | ✅ | 8K | $0.20 / $2.00 | Vision + short context |
moonshot-v1-32k-vision-preview | ❌ | ✅ | 32K | $1.00 / $3.00 | Vision + medium context |
moonshot-v1-128k-vision-preview | ❌ | ✅ | 128K | $2.00 / $5.00 | Vision + long context |
Prices: Per 1M tokens. Cache pricing applies to prompt tokens served from automatic context cache (only Kimi K2.x models support cache).
CRITICAL — Kimi K2.6/K2.5 parameter constraints: API returns
invalid_request_errorfor any deviation:
temperature: MUST be1.0in thinking mode, MUST be0.6in non-thinking modetop_p: MUST be0.95n: MUST be1presence_penaltyandfrequency_penalty: MUST be0(not modifiable)Moonshot V1 models use standard OpenAI-compatible parameters with no such constraints.
Default Agent Configuration:
Strategy: prefer kimi-k2.5 as cost-effective workhorse (36% cheaper input vs kimi-k2.6); reserve kimi-k2.6 for critical reasoning. All kimi-k2.x agents are configured with the API-required fixed parameters (temp/top_p/n) and explicit extra_body.thinking.type. For thinking-enabled agents, extra_body.thinking.keep: "all" is set to preserve historical reasoning_content in multi-turn tool call chains (without it Moonshot returns "thinking is enabled but reasoning_content is missing").
| Agent Role | Default Model | Thinking | Temperature | Top P | Max Output |
|---|---|---|---|---|---|
| Generator / Refiner | kimi-k2.6 | Enabled (keep=all) | 1.0 | 0.95 | 32768 |
| Coder | kimi-k2.6 | Enabled (keep=all) | 1.0 | 0.95 | 20480 |
| Pentester | kimi-k2.6 | Enabled (keep=all) | 1.0 | 0.95 | 16384 |
| Adviser (mentor/planner) | kimi-k2.6 | Enabled (keep=all) | 1.0 | 0.95 | 8192 |
| Primary Agent / Assistant | kimi-k2.5 | Enabled (keep=all) | 1.0 | 0.95 | 16384 |
| Installer | kimi-k2.5 | Enabled (keep=all) | 1.0 | 0.95 | 12288 |
| Reflector / Searcher / Enricher | kimi-k2.5 | Disabled | 0.6 | 0.95 | 4096 |
| Simple / Simple JSON | kimi-k2.5 | Disabled | 0.6 | 0.95 | 2048 |
Key Features:
extra_body.thinking.typethinking.keep: "all" preserves historical reasoning_content across turns — required for multi-turn tool call chainsMulti-turn with thinking + tool calls: PentAGI's universal reasoning preservation pattern (TextPartWithReasoning + WithPreserveReasoningContent) automatically ensures reasoning_content is sent back in the required TextContent → ToolCall order, satisfying Moonshot's "thinking is enabled but reasoning_content is missing in assistant tool call message" requirement.
LiteLLM Integration: Set KIMI_PROVIDER=moonshot to enable model name prefixing when using default PentAGI configurations with LiteLLM proxy. Leave empty for direct API usage.
PentAGI integrates with Qwen from Alibaba Cloud Model Studio (DashScope), providing powerful multilingual models with reasoning capabilities and context caching support.
| Variable | Default Value | Description |
|---|---|---|
QWEN_API_KEY | Qwen API key for authentication | |
QWEN_SERVER_URL | https://dashscope-us.aliyuncs.com/compatible-mode/v1 | Qwen API endpoint URL (international) |
QWEN_PROVIDER | Provider prefix for LiteLLM integration (optional) |
# Direct API usage (Global/US endpoint)
QWEN_API_KEY=your_qwen_api_key
QWEN_SERVER_URL=https://dashscope-us.aliyuncs.com/compatible-mode/v1
# Alternative endpoints
QWEN_SERVER_URL=https://dashscope-intl.aliyuncs.com/compatible-mode/v1 # International (Singapore)
QWEN_SERVER_URL=https://dashscope.aliyuncs.com/compatible-mode/v1 # Chinese Mainland (Beijing)
# With LiteLLM proxy
QWEN_API_KEY=your_litellm_key
QWEN_SERVER_URL=http://litellm-proxy:4000
QWEN_PROVIDER=dashscope # Adds prefix to model names (dashscope/qwen-plus) for LiteLLM
PentAGI supports 33 Qwen models curated for agent workflows: text reasoning, code generation, and vision-language (browser screenshots). All models are non-snapshot main aliases with tool calling, streaming, thinking modes, and context caching. Models marked with * are used in default configuration.
Flagship Models (Top-tier Reasoning)
| Model ID | Thinking | Intl | Global/US | China | Price (Input/Output/Cache) | Use Case |
|---|---|---|---|---|---|---|
qwen3.7-max* | ✅ | ✅ | ✅ | ✅ | $2.50/$7.50/$0.50 | Next-gen flagship for agent-centric era (generator/refiner/adviser default) |
qwen3.6-max-preview | ✅ | ✅ | ✅ | ✅ | $1.30/$7.80/$0.13 | Preview Max with enhanced vibe coding & front-end skills |
qwen3-max | ✅ | ✅ | ✅ | ✅ | $1.20/$6.00/$0.24 | Previous-gen flagship with agent programming upgrades |
qwen-plus | ✅ | ✅ | ✅ | ✅ | $0.40/$4.00/$0.08 | Qwen3-backbone Plus with switchable thinking modes |
Balanced Plus Models (Mid-tier)
| Model ID | Thinking | Intl | Global/US | China | Price (Input/Output/Cache) | Use Case |
|---|---|---|---|---|---|---|
qwen3.6-plus* | ✅ | ✅ | ✅ | ✅ | $0.50/$3.00/$0.05 | Native VL Plus with agentic coding (primary/assistant/pentester default) |
qwen3.5-plus | ✅ | ✅ | ✅ | ✅ | $0.40/$2.40/$0.04 | Previous-gen native VL with strong multimodal capabilities |
Fast Flash Models (Cost-optimized)
| Model ID | Thinking | Intl | Global/US | China | Price (Input/Output/Cache) | Use Case |
|---|---|---|---|---|---|---|
qwen3.6-flash | ✅ | ✅ | ✅ | ✅ | $0.25/$1.50/$0.025 | Latest Flash with significant agentic-coding boost |
qwen3.5-flash* | ✅ | ✅ | ✅ | ✅ | $0.10/$0.40/$0.01 | Ultra-fast lightweight (simple/reflector/searcher/enricher default) |
qwen-flash | ✅ | ✅ | ✅ | ✅ | $0.05/$0.40/$0.01 | Qwen3-series Flash with 1M context, tiered pricing |
Code-Specialized Models
| Model ID | Thinking | Intl | Global/US | China | Price (Input/Output/Cache) | Use Case |
|---|---|---|---|---|---|---|
qwen3-coder-plus* | ❌ | ✅ | ✅ | ✅ | $1.00/$5.00/$0.20 | Strong coding agent with autonomous programming (coder default) |
qwen3-coder-flash* | ❌ | ✅ | ✅ | ✅ | $0.30/$1.50/$0.06 | Fast code-gen with multi-turn tool stability (installer default) |
qwen3-coder-next | ❌ | ✅ | ✅ | ✅ | $0.30/$1.50/— | Open-source code generation, SOTA at same scale |
Vision-Language Models (Browser & Screenshot Analysis)
| Model ID | Thinking | Intl | Global/US | China | Price (Input/Output/Cache) | Use Case |
|---|---|---|---|---|---|---|
qwen3-vl-plus | ✅ | ✅ | ✅ | ✅ | $0.20/$1.60/$0.04 | VL with visual agent capabilities, ultra-long video understanding |
qwen3-vl-flash | ✅ | ✅ | ✅ | ✅ | $0.05/$0.40/$0.01 | Small VL with 2D/3D localization for browser triage |
qvq-max | ✅ | ✅ | ✅ | ✅ | $1.20/$4.80/— | Visual reasoning with chain-of-thought |
Open-Source Qwen3.6 Series
| Model ID | Thinking | Intl | Global/US | China | Price (Input/Output/Cache) | Use Case |
|---|---|---|---|---|---|---|
qwen3.6-27b | ✅ | ✅ | ✅ | ✅ | $0.60/$3.60/— | Native VL on hybrid architecture, on-premises ready |
qwen3.6-35b-a3b | ✅ | ✅ | ✅ | ✅ | $0.25/$1.49/— | Efficient 35B MoE (~3B active) for continuous monitoring |
Open-Source Qwen3.5 Series
| Model ID | Thinking | Intl | Global/US | China | Price (Input/Output/Cache) | Use Case |
|---|
| Parameter | Environment Variable | Default | Description |
|---|
| Preserve Last | ASSISTANT_SUMMARIZER_PRESERVE_LAST | true | Whether to preserve all messages in the assistant's last section |
| Last Section Size | ASSISTANT_SUMMARIZER_LAST_SEC_BYTES | 76800 | Maximum byte size for assistant's last section (75KB) |
| Max Body Pair Size | ASSISTANT_SUMMARIZER_MAX_BP_BYTES | 16384 | Maximum byte size for a single body pair in assistant context (16KB) |
| Max QA Sections | ASSISTANT_SUMMARIZER_MAX_QA_SECTIONS | 7 | Maximum QA sections to preserve in assistant context |
| Max QA Size | ASSISTANT_SUMMARIZER_MAX_QA_BYTES | 76800 | Maximum byte size for assistant's QA sections (75KB) |
| Keep QA Sections | ASSISTANT_SUMMARIZER_KEEP_QA_SECTIONS | 3 | Number of recent QA sections to preserve without summarization |
cohere.command-r-plus-v1:0| Cohere |
| ❌ |
| ❌ |
| $3.00/$15.00 |
| Large-scale operations, superior RAG |
deepseek.v3.2 | DeepSeek | ❌ | ❌ | $0.58/$1.68 | Long-context reasoning, efficiency |
openai.gpt-oss-120b-1:0* | OpenAI (OSS) | ✅ | ❌ | $0.15/$0.60 | Strong reasoning, scientific analysis |
openai.gpt-oss-20b-1:0 | OpenAI (OSS) | ✅ | ❌ | $0.07/$0.30 | Efficient coding, software development |
qwen.qwen3-next-80b-a3b | Qwen | ❌ | ❌ | $0.15/$1.20 | Ultra-long context, flagship reasoning |
qwen.qwen3-32b-v1:0 | Qwen | ❌ | ❌ | $0.15/$0.60 | Balanced reasoning, research use cases |
qwen.qwen3-coder-30b-a3b-v1:0 | Qwen | ❌ | ❌ | $0.15/$0.60 | Vibe coding, natural-language first |
qwen.qwen3-coder-next | Qwen | ❌ | ❌ | $0.45/$1.80 | Tool use, function calling optimized |
mistral.mistral-large-3-675b-instruct | Mistral | ❌ | ✅ | $4.00/$12.00 | Advanced multimodal, long-context |
moonshotai.kimi-k2.5 | Moonshot | ❌ | ✅ | $0.60/$3.00 | Vision, language, code in one model |