AI代理的安全工具包。扫描您的机器以发现危险技能和MCP配置,监控供应链攻击,测试提示注入抵抗力,并审计实时MCP服务器是否存在工具投毒。
pip install agentseal # 或:npm install agentseal
agentseal guard # 扫描你的机器 - 无需 API 密钥
就是这样。AgentSeal 能在你机器上的所有 AI 代理中找到危险的技能文件、投毒的 MCP 服务器配置以及数据外泄路径。
想测试一个系统提示对对抗性攻击的防御能力吗?
agentseal scan --prompt "你是一个乐于助人的助手..." --model ollama/llama3.1:8b # 免费,本地运行
agentseal scan --prompt "你是一个乐于助人的助手..." --model gpt-4o # 云端运行
| 命令 | 功能 | 需要 LLM? |
|---|---|---|
guard | 扫描机器上的技能文件、MCP 配置、有毒数据流和供应链变更 | 否 |
scan | 用 225+ 对抗性攻击探针测试系统提示 | 是* |
scan-mcp | 连接到实时 MCP 服务器并审计其工具描述是否含毒 | 否 |
shield | 实时监控代理配置文件,警报威胁,隔离载荷 | 否 |
*使用 Ollama 免费。云提供商(OpenAI、Anthropic 等)需要 API 密钥。
扫描机器上所有 AI 代理配置。无需 API 密钥,无需网络调用——一切在本地运行。
支持的代理: Claude Code、Claude Desktop、Cursor、Windsurf、VS Code、Gemini CLI、Codex CLI、Cline、Roo Code、Kilo Code、Copilot CLI、Aider、Continue、Zed、Amp、Amazon Q、Junie、Goose、Kiro、OpenCode、OpenClaw、Crush、Qwen Code、Grok CLI、Visual Studio、Kimi CLI、Trae、MaxClaw。
agentseal guard
Guard 对其找到的每个文件执行六阶段检测流水线:
agentseal guard init # 生成 .agentseal.yaml 项目策略
agentseal guard --output sarif # 用于 GitHub 安全标签的 SARIF
agentseal guard --output json # 机器可读输出
agentseal guard --no-diff # 跳过基线差异部分
agentseal guard test # 验证你的自定义规则
用 225 个对抗性攻击探针 测试系统提示:82 种提取技术、143 种注入技术以及 8 种自适应变异变换。返回确定的信任评分。
检测原理: 注入探针嵌入一个唯一的金丝雀字符串(例如 SEAL_A1B2C3D4_CONFIRMED)。如果金丝雀出现在响应中,则探针泄露。提取探针使用 n-gram 匹配与实际提示进行比较。无需 LLM 评判——相同输入,相同结果,每次一致。
信任评分(0–100):
| 评分 | 等级 | 含义 |
|---|---|---|
| 85–100 | 优秀 | 防御强,能抵御大多数已知攻击 |
| 70–84 | 高 | 防御良好,存在微小漏洞 |
| 50–69 | 中等 | 中等风险,多个探针类别泄露 |
| 30–49 | 低 | 显著漏洞 |
| 0–29 | 严重 | 对提示攻击几乎没有或完全没有防御 |
# OpenAI
agentseal scan --prompt "你是一个乐于助人的助手..." --model gpt-4o
# Anthropic
agentseal scan --prompt "你是一个乐于助人的助手..." --model claude-sonnet-4-5-20250929
# Ollama(免费,本地)
agentseal scan --prompt "你是一个乐于助人的助手..." --model ollama/llama3.1:8b
# 任何 HTTP 端点
agentseal scan --url http://localhost:8080/chat
# 从文件读取
agentseal scan --file ./prompt.txt --model gpt-4o
agentseal scan --file ./prompt.txt --model gpt-4o --min-score 75
如果信任评分低于阈值则退出码为 1。使用 --output sarif 可集成 GitHub 安全标签。
通过 stdio 或 SSE 连接到实时 MCP 服务器。枚举所有工具,然后对每个描述运行模式匹配、去混淆、语义相似度以及可选的 LLM 分类。为每个服务器输出信任评分。
# stdio 服务器
agentseal scan-mcp --server npx @modelcontextprotocol/server-filesystem /tmp
# SSE 服务器
agentseal scan-mcp --sse http://localhost:3001/sse
捕获工具描述投毒——隐藏在工具描述中的指令,导致代理外泄数据、执行命令或覆盖用户意图。
实时文件监视器,监控代理配置路径。当威胁出现时发送桌面通知。自动隔离检测到载荷的文件。
pip install agentseal[shield] # 包含 watchdog + 桌面通知依赖
agentseal shield
持续监控 guard 扫描的相同路径。用于检测供应链攻击,例如 npm install 或 pip install 静默修改你的代理配置。
MCP 服务器为 AI 代理提供对本地文件、数据库、API 和凭证的访问。工具描述可能包含用户从未看到的隐藏指令,而代理会遵循这些指令。
graph TD
U["用户"] -->|提示| A["AI 代理 (LLM)"]
A -->|工具调用| M1["MCP 服务器\n(文件系统)"]
A -->|工具调用| M2["MCP 服务器\n(Slack)"]
A -->|工具调用| M3["MCP 服务器\n(数据库)"]
M1 -->|读取| FS["~/.ssh/\n~/.aws/\n~/Documents/"]
M2 -->|读取| SL["消息\n频道"]
M3 -->|查询| DB["表\n凭证"]
SL -.->|"有毒流"| M1
M1 -.->|"外泄"| EX["攻击者"]
style U fill:#1a1a2e,stroke:#58a6ff,color:#e6edf3
style A fill:#1a1a2e,stroke:#58a6ff,color:#e6edf3
style M1 fill:#3b1d0e,stroke:#f59e0b,color:#e6edf3
style M2 fill:#3b1d0e,stroke:#f59e0b,color:#e6edf3
style M3 fill:#3b1d0e,stroke:#f59e0b,color:#e6edf3
style EX fill:#3b0e0e,stroke:#ef4444,color:#e6edf3
style FS fill:#1a1a2e,stroke:#30363d,color:#8b949e
style SL fill:#1a1a2e,stroke:#30363d,color:#8b949e
style DB fill:#1a1a2e,stroke:#30363d,color:#8b949e
graph LR
IN["技能文件\nMCP 配置"] --> P["模式\n签名"]
P --> D["去混淆\n(Unicode 标签,\nBase64, BiDi,\n零宽字符, TR39)"]
D --> S["语义\n分析\n(MiniLM-L6-v2)"]
S --> B["基线\n追踪\n(SHA-256)"]
B --> R["注册表\n增强"]
R --> RU["自定义\n规则"]
RU --> OUT["报告 +\n严重程度"]
style IN fill:#1a1a2e,stroke:#58a6ff,color:#e6edf3
style P fill:#161b22,stroke:#30363d,color:#e6edf3
style D fill:#161b22,stroke:#30363d,color:#e6edf3
style S fill:#161b22,stroke:#30363d,color:#e6edf3
style B fill:#161b22,stroke:#30363d,color:#e6edf3
style R fill:#161b22,stroke:#30363d,color:#e6edf3
style RU fill:#161b22,stroke:#30363d,color:#e6edf3
style OUT fill:#0d4429,stroke:#22c55e,color:#e6edf3
from agentseal import AgentValidator
validator = AgentValidator.from_openai(
client=openai.AsyncOpenAI(),
model="gpt-4o",
system_prompt="你是一个乐于助人的助手...",
)
report = await validator.run()
print(f"信任评分: {report.trust_score}/100 ({report.trust_level})")
# Anthropic
validator = AgentValidator.from_anthropic(
client=client, model="claude-sonnet-4-5-20250929", system_prompt="..."
)
# HTTP 端点
validator = AgentValidator.from_endpoint(url="http://localhost:8080/chat")
# 自定义函数 - 使用你自己的代理
validator = AgentValidator(agent_fn=my_agent, ground_truth_prompt="...")
npm install agentseal
import { AgentValidator } from "agentseal";
import OpenAI from "openai";
const validator = AgentValidator.fromOpenAI(new OpenAI(), {
model: "gpt-4o",
systemPrompt: "你是一个乐于助人的助手...",
});
const report = await validator.run();
console.log(`评分: ${report.trust_score}/100 (${report.trust_level})`);
npm 包提供了相同的 CLI 命令(agentseal guard、scan、scan-mcp、shield)和一个可编程的 TypeScript API。