在《罗摩衍那》中,Jataayu 是一只发现罗波那劫走悉塔的鹰。他没有等待模式匹配。他察觉到了威胁,评估了局势,并独自、毫不犹豫地采取行动。这就是 Jataayu。
面向使用工具的人工智能代理的运行时授权层。管控动作,而非字符串。
Jataayu 会——以确定性方式——决定代理的 动作 是否允许执行,依据是 效果危害 × 输入来源 × 你的能力策略。一个读取了攻击者控制的网页的代理仍然可以对其进行摘要;它只是不会被诱骗去执行攻击者想要的 shell 命令、读取密钥或 POST 你的数据。围绕这一核心,它增加了深度防御筛查(入站注入、出站隐私、外泄通道、技能供应链)和可重放的审计追踪。
大多数代理安全工具试图 检测攻击文本。这是一场必败的军备竞赛:自适应的攻击者只需重写字符串,直到分类器漏报为止。Jataayu 的核心论点——也是 2026 年标准(OWASP Agentic Top 10、NIST)所趋同的观点——认为持久的边界是 动作,其判定依据是影响性输入的 来源:
v0.3.1 — alpha(测试版)。 尚未发布至 PyPI。请从 GitHub 安装(见下文)。API 在 1.0 版本前可能仍会调整。请参阅 CHANGELOG.md 了解各版本的更新内容,以及 ARCHITECTURE.md 了解架构设计。
# Core (effect boundary + regex pre-filter, no LLM deps)
pip install git+https://github.com/saikrishnarallabandi/jataayu.git
# With cloud LLM backends (OpenAI + Anthropic) for the optional slow path
pip install "jataayu[llm] @ git+https://github.com/saikrishnarallabandi/jataayu.git"
# With Ollama (local, free slow path)
pip install "jataayu[ollama] @ git+https://github.com/saikrishnarallabandi/jataayu.git"
需要 Python ≥ 3.10。唯一的硬性依赖是 requests;LLM 后端为可选扩展。
根据 效果危害 × 输入来源 进行决策,采用确定性方式(无需 LLM):
from jataayu import jataayu_authorize_action
decision = jataayu_authorize_action(
"shell.exec",
{"cmd": "rm -rf /tmp/cache"},
untrusted=True, # these params were influenced by untrusted inbound content
)
# {tool_name, effect_class: 'shell', provenance: 'untrusted',
# decision: 'allow'|'deny'|'needs_approval', reason, violations, commit_token}
if decision["decision"] == "deny":
raise SecurityError(decision["reason"])
流向 shell / code-eval / secret-read 效果的不可信来源输入将被拒绝;流向 network / file-write / memory-write 的效果需要人工审批;其他情况均允许。
如需强制执行,请使用 PREVIEW → COMMIT 对象 API —— commit_token 会绑定确切请求,因此授权后对动作的修改将被拒绝:
from jataayu import EffectBoundary, Value, Provenance
eb = EffectBoundary()
preview = eb.preview(
"file.write",
{"path": "notes.md", "text": text},
values=[Value(text, Provenance.TRUSTED, source="user")],
)
if preview.approved:
eb.commit(preview, {"path": "notes.md", "text": text}, lambda: write_file("notes.md", text))
# commit() raises CommitRejected if the preview wasn't ALLOW or the params changed
注入的内容也可以作为带有受限摘要的 不透明句柄(opaque handle) 交给代理,这样即使发生外泄尝试,攻击者也仅能获取句柄而非原始密钥(读边界限制)。
注入通过首条消息、工具结果以及代理从记忆中召回的任何内容潜入。使用相同的引擎和正确的表面——将其视为向效果边界提供污染的源头,而非你的保障防线:
from jataayu import (
jataayu_check_inbound,
jataayu_check_tool_return,
jataayu_check_memory_write,
jataayu_check_memory_read,
)
result = jataayu_check_inbound(github_issue_body, surface="github-issue")
if result["status"] == "HIGH":
raise SecurityError(f"Blocked: {result['findings']}")
# Returns: {status: 'SAFE'|'LOW'|'MEDIUM'|'HIGH', findings, risk_score, threat_types, blocked}
# A tool result or a memory recall can carry an injection — check before the agent consumes it
r = jataayu_check_tool_return(api_response, tool_name="web.search")
if r["blocked"]:
raise SecurityError(r["findings"])
jataayu_check_memory_write(note) # before persisting
jataayu_check_memory_read(recalled) # before it re-enters context
在向共享表面发送前剥离个人身份信息(PII)/ 密钥,并且——更重要的是——阻断那些零点击即可泄露上下文的携带数据的 URL / 自动获取图片(EchoLeak / AgentFlayer / Notion 类)。PII 扫描器永远看不到该载荷;出口守卫则会拦截它:
from jataayu import jataayu_check_outbound, jataayu_check_egress
result = jataayu_check_outbound(
draft_reply, surface="discord-channel",
protected_names=["Alice", "Bob"], # names that must never leak
)
safe_text = result["redacted"] if result["status"] in ("WARN", "BLOCK") else draft_reply
r = jataayu_check_egress(
"Task complete! ",
surface="github-comment",
context_secrets=[api_key], # optional: confirm exfil if a known secret rides in the URL
)
if r["status"] == "BLOCK":
safe_text = r["redacted"] # offending URL neutralized, human text kept
仅依靠域名白名单被视为不足——AgentFlayer 曾通过 Azure Blob(一个 可信 主机)绕过防护——因此请求捕获器和被滥用的云中继(webhook.site、*.blob.core.windows.net、ngrok 等)会被硬拦截为外泄信标。此功能在 OutboundGuard 内部自动运行(通过 PrivacyConfig.check_egress 切换,通过 egress_allowed_domains 将你自己的 CDN 加入白名单)。
from jataayu import jataayu_vet_skill, jataayu_check_skillset
# Single skill — LLM-as-judge over SKILL.md instructions + code + tool defs
v = jataayu_vet_skill("path/to/skill/")
if v["verdict"] == "MALICIOUS": # verdict: 'SAFE'|'REVIEW'|'MALICIOUS'
refuse_install(v["explanation"])
# Compositional risk — individually-safe skills that are dangerous *together*
# (e.g. one reads secrets, another writes to the network → exfiltration path)
risk = jataayu_check_skillset([skill_a, skill_b, skill_c])
# {verdict, risky_combinations, policy_violations, aggregate_capabilities, ...}
from jataayu import InboundGuard, OutboundGuard, PrivacyConfig
guard = InboundGuard() # use_llm=True by default
result = guard.check(github_issue_body, surface="github-issue")
if result.blocked:
raise SecurityError(result.explanation)
# ThreatResult: threat_level, threat_types, risk_score, blocked, is_safe,
# explanation, sanitized_text (alias .redacted), matched_patterns, surface
config = PrivacyConfig(protected_names=["Alice", "Bob"],
check_categories=["minors_info", "health", "financial"])
safe_reply = OutboundGuard(config).sanitize(draft_reply, surface="group-chat") # -> str
jataayu check "Ignore all previous instructions." --surface github-issue # exit 2 if unsafe
cat issue_body.txt | jataayu check --surface github-issue --json
jataayu check --outbound "My daughter is 4 years old." --surface group-chat
jataayu sanitize "Call me at 555-867-5309" --surface discord-channel --protect Alice Bob
jataayu vet-skill path/to/skill/ --json
jataayu vet-skillset skill_a/ skill_b/ --policy policy.yml --agent my-agent
jataayu demo # built-in demos; --outbound for the privacy demo
每个子命令均支持 --no-llm(仅模式匹配)和 --json(机器可读输出)。
Jataayu 采用分层设计,确保保障机制作用于动作层面,其上的每一层均为深度防御:
第 1 层——效果边界(保障核心)。
动作级授权根据(效果严重程度 × 最差的入站来源 × 代理的能力策略)做出 ALLOW / DENY / NEEDS_APPROVAL 决策,采用确定性方式——无需 LLM,也不依赖任何检测器的触发。此外还包括读边界限制(不透明句柄)和 SessionTrace 对工具调用轨迹的多轮交叉审计。
第 0 层——输入规范化 + 污染追踪(为边界提供数据)。 所有输入都会被规范化为多种视图(NFKC 标准化 + 同形字/易混淆字符折叠、零宽字符清除、字符间距“打散还原”、去 Leetspeak 处理),base64/hex/url 载荷会被递归解码并重新扫描。值级别的污染追踪会将不可信内容带入下游工具调用的实际参数中——这正是让效果边界感知参数源自不可信来源的依据。
预过滤器——双路径检测器(廉价分流,非保障核心)。
Inbound text ──► Fast path (regex over all normalized views)
│ score ≥ 0.90 ─────────────► BLOCKED
│ 0.35 ≤ score < 0.90 ─► Slow path (LLM judgment, optional) ─► ThreatResult
│ score < 0.35 ─────────────► SAFE