Google Cloud Hackathon 提交项目 — 由 Google Cloud Agent Builder (Vertex AI Agent) 编排的自主威胁检测与响应,基于 Gemini 2.0 Flash、模型上下文协议(MCP)和 Dynatrace 可观测性。
安全团队在处理典型安全事件时,平均检测时间(MTTD)为 30分钟,平均修复时间(MTTR)为 4小时。等到人工分析师检测到攻击、编写防火墙规则并部署完毕,损失早已造成。
SentinelMCP 将 MTTD 缩短至 < 30秒,MTTR 缩短至 < 30秒。它部署了一个 Google Cloud Agent Builder 代理,接收 Dynatrace 异常 Webhook,通过 Gemini 2.0 Flash 进行推理,然后通过 MCP 调用真实的防御工具——直接在实时基础设施上执行操作,而不仅仅是发出告警。
Dynatrace webhook → Agent Builder 接收告警 → Gemini 推理 → MCP 工具执行 → 威胁消除
1 秒 2 秒 8 秒 15 秒 总计 25 秒
最终结果:过去需要人工分析师耗时 4 小时的事件生命周期,现在可在 30 秒内自主完成,并在 MongoDB 中保留完整的审计轨迹,通过 Jaeger 实现分布式链路追踪。
┌──────────────────────────────────────────────────────────────────────┐
│ 沙箱环境 │
│ │
│ ┌─────────────┐ 100–1000 req/s ┌──────────────────────────┐ │
│ │ 攻击方 │ ──────────────────►│ 受害者服务 │ │
│ │ (DDoS/SQLi)│ │ 端口 8080 │ │
│ └─────────────┘ │ /admin/* 控制 API │ │
│ └────────────┬─────────────┘ │
└───────────────────────────────────────────────────│──────────────────┘
│ 指标 + 异常 webhook
┌────────────────────────────────▼─────────────────┐
│ Dynatrace Mock (异常引擎) │
│ 威胁风险评分 → 触发 POST /alerts/simulate │
└────────────────────────┬─────────────────────────┘
│ webhook
┌─────────────────────────────────▼──────────────────────────┐
│ Google Cloud Agent Builder (Vertex AI Agent) │
│ ┌──────────────────────────────────────────────────────┐ │
│ │ Gemini 2.0 Flash ── 对威胁上下文进行推理 │ │
│ │ 工具清单 ── 从 MCP 动态获取 │ │
│ │ HITL 门控 ── 置信度低于阈值时暂停 │ │
│ └──────────────────────────────────────────────────────┘ │
└──────────────┬─────────────────────────┬───────────────────┘
│ REST 调用 │ 审计写入
┌───────────────▼──────┐ ┌────────────▼──────────┐
│ MCP 服务器 :8001 │ │ MongoDB 7.0 Atlas │
│ 11 个真实工具 │ │ 事件 + 轨迹 │
└───────────┬──────────┘ └───────────────────────┘
│ 作用于实时基础设施
┌───────────▼──────────┐
│ FastAPI :8000 │
│ victims-service :8080 │
└──────────────────────┘
┌──────────────────────────────────────────────────┐
│ 可观测性栈 │
│ OTel Collector ──► Jaeger (分布式链路追踪) │
│ Prometheus ──────────────► Grafana │
└──────────────────────────────────────────────────┘
关键设计原则: 优雅降级 + 人在回路(HITL)。Agent Builder 代理在无法以足够置信度做出决策时,会安全暂停并将控制权交还给操作员——它永远不会在关键基础设施上盲目猜测。
三个服务均已部署并运行在 Google Cloud Run 上:
针对生产环境发起实时攻击:
.\attack.production.ps1 status # 验证所有服务健康
.\attack.production.ps1 ddos # DDoS — Gemini 封锁 IP + 限速
.\attack.production.ps1 sql # SQL 注入 — Gemini 激活 WAF
.\attack.production.ps1 brute # 暴力破解 — Gemini 封锁源 IP
.\attack.production.ps1 stop # 重置防御
攻击容器在本地运行,通过 HTTPS 攻击 Cloud Run 上的 victim-service。在指挥中心仪表盘实时观察检测→推理→消除的完整周期。
Google Cloud Agent Builder (Vertex AI ReasoningEngine) 是 SentinelMCP 的核心编排器。实现代码位于 sentinel_mcp/agent/agent_builder.py。
SentinelAgent 继承自 vertexai.preview.reasoning_engines.Queryable——这是 Agent Builder 的官方编程接口——可以在本地运行,也可以通过一条命令部署到托管的 Agent Builder 服务中。
| 职责 | 实现方式 |
|---|---|
| Webhook 接收 | SentinelAgent.query(incident=...) — Agent Builder 在每次 Dynatrace 告警时调用 |
| 工具声明 | set_up() 中将 6 个 MCP 工具注册为 Vertex AI FunctionDeclaration 对象 |
| 推理 | Gemini 2.0 Flash 读取威胁上下文 + 工具目录,选择最小工具集,自动调用 |
| HITL 门控 | 如果 Gemini 未返回任何工具调用(置信度低于阈值),则触发 escalate_to_humans=True |
| 审计 | 每次工具调用记录到 MongoDB Atlas,并作为 OTel span 导出到 Jaeger |
关键代码 — sentinel_mcp/agent/agent_builder.py:
class SentinelAgent(reasoning_engines.Queryable):
def set_up(self):
# MCP 工具作为 Vertex AI FunctionDeclarations — Gemini 运行时选择
mcp_tools = Tool(function_declarations=[
FunctionDeclaration(name="block_ip_address", ...),
FunctionDeclaration(name="activate_waf", ...),
FunctionDeclaration(name="rate_limit_requests", ...),
FunctionDeclaration(name="scale_service", ...),
FunctionDeclaration(name="collect_forensic_logs", ...),
FunctionDeclaration(name="analyze_attack_pattern", ...),
])
self._model = GenerativeModel("gemini-2.0-flash-001", tools=[mcp_tools])
self._chat = self._model.start_chat()
def query(self, *, incident: dict) -> dict:
# 入口 — Agent Builder 每次收到 Dynatrace webhook 时调用
response = self._chat.send_message(threat_prompt(incident))
tool_calls = [{"name": p.function_call.name, "args": dict(p.function_call.args)}
for p in response.candidates[0].content.parts if p.function_call]
return {"engine": "agent_builder", "tool_calls": tool_calls, ...}
部署到 Google Cloud Agent Builder:
# 本地测试(使用 google_ai_studio 后端无需 GCP)
python scripts/deploy_agent_builder.py --test-only
# 部署到托管的 Agent Builder 端点
python scripts/deploy_agent_builder.py --project my-gcp-project
将 Agent Builder 设为活动后端:
# .env
GEMINI_BACKEND=agent_builder
GOOGLE_CLOUD_PROJECT=my-gcp-project
VERTEX_LOCATION=us-central1
SentinelMCP 在三个层面集成 Dynatrace 可观测性:
OpenTelemetry 管道 — sentinel-otel-collector 从 FastAPI 核心采集 span 和指标,并将其转发到 Dynatrace Mock 端点(OTLP/HTTP :14318)。每个事件、工具执行和 LLM 调用都通过完整关联 ID 进行端到端链路追踪。
带自动触发的异常引擎 — dynatrace-mock 运行后台循环,监控 http_requests_total,计算威胁风险评分(dynatrace_mock_sentinel_risk_score,0–10),并在评分超过关键阈值时自动向 SentinelMCP 发送 POST /api/v1/alerts/simulate/{type}——复现生产环境中 Dynatrace Davis 异常检测 webhook 的精确行为。
由 Dynatrace 指标驱动的 Grafana 仪表盘 — 风险评分仪表和趋势图直接使用 dynatrace_mock_sentinel_risk_score,可以直观地看到 Dynatrace 标记异常的确切时刻与 SentinelMCP 完成消除的时刻。
Gemini 代理在运行时动态发现其防御能力——无需硬编码工具列表,功能变更时无需重写提示:
GET /mcp/tools → 返回 11 个工具描述 → Gemini 选择并调用工具
每个工具在实时目录中标记为 REAL 或 Audit-Only。Gemini 读取此标志,了解自己能在实时基础设施上做什么,并在推理轨迹中解释每个工具选择。增加新的防御能力无需任何提示工程。
作用于实时基础设施的工具: