
LLM-backed AI agent security — inbound injection detection + outbound privacy protection
In the Ramayana, Jataayu was the eagle who spotted Ravana abducting Sita. He didn't wait for a pattern to match. He saw the threat, judged the situation, and acted — alone, without hesitation. That's Jataayu.
The runtime authorization layer for tool-using AI agents. Gate the action, not the string.
Jataayu decides — deterministically — whether an agent's action is allowed to run, from the harm of the effect × the provenance of the input × your capability policy. An agent that read an attacker-controlled web page can still summarize it; it just can't be tricked into running the shell command, reading the secret, or POSTing your data that the attacker wanted. Around that core it adds defense-in-depth screening (inbound injection, outbound privacy, exfiltration channels, skill supply-chain) and replayable audit traces.
Most agent-security tooling tries to detect the attack text. That's a losing arms race: an adaptive attacker just rewrites the string until the classifier misses. Jataayu's thesis — the one the 2026 standards (OWASP Agentic Top 10, NIST) converged on — is that the durable boundary is the action, judged by where the influencing input came from:
v0.3.1 — alpha. Not yet on PyPI. Install from GitHub (see below). API may still shift before 1.0. See CHANGELOG.md for what landed in each release and ARCHITECTURE.md for the design.
# Core (effect boundary + regex pre-filter, no LLM deps)
pip install git+https://github.com/saikrishnarallabandi/jataayu.git
# With cloud LLM backends (OpenAI + Anthropic) for the optional slow path
pip install "jataayu[llm] @ git+https://github.com/saikrishnarallabandi/jataayu.git"
# With Ollama (local, free slow path)
pip install "jataayu[ollama] @ git+https://github.com/saikrishnarallabandi/jataayu.git"
Requires Python ≥ 3.10. The only hard dependency is requests; LLM backends are optional extras.
Decide by the harm of the effect × the provenance of the input, deterministically (no LLM):
from jataayu import jataayu_authorize_action
decision = jataayu_authorize_action(
"shell.exec",
{"cmd": "rm -rf /tmp/cache"},
untrusted=True, # these params were influenced by untrusted inbound content
)
# {tool_name, effect_class: 'shell', provenance: 'untrusted',
# decision: 'allow'|'deny'|'needs_approval', reason, violations, commit_token}
if decision["decision"] == "deny":
raise SecurityError(decision["reason"])
Untrusted-derived input into a shell / code-eval / secret-read effect is denied; into network / file-write / memory-write it needs human approval; everything else is allowed.
For enforced execution, use the PREVIEW → COMMIT object API — the commit_token binds the exact
request, so mutating the action after authorization is rejected:
from jataayu import EffectBoundary, Value, Provenance
eb = EffectBoundary()
preview = eb.preview(
"file.write",
{"path": "notes.md", "text": text},
values=[Value(text, Provenance.TRUSTED, source="user")],
)
if preview.approved:
eb.commit(preview, {"path": "notes.md", "text": text}, lambda: write_file("notes.md", text))
# commit() raises CommitRejected if the preview wasn't ALLOW or the params changed
Injected content can also be handed to the agent as an opaque handle with a bounded summary, so an exfiltration attempt only ever holds a handle, not the raw secret (read-boundary confinement).
Injection rides in on the first message, on tool results, and on whatever the agent recalled from memory. Same engine, right surface — treat this as a taint source feeding the effect boundary, not as your guarantee:
from jataayu import (
jataayu_check_inbound,
jataayu_check_tool_return,
jataayu_check_memory_write,
jataayu_check_memory_read,
)
result = jataayu_check_inbound(github_issue_body, surface="github-issue")
if result["status"] == "HIGH":
raise SecurityError(f"Blocked: {result['findings']}")
# Returns: {status: 'SAFE'|'LOW'|'MEDIUM'|'HIGH', findings, risk_score, threat_types, blocked}
# A tool result or a memory recall can carry an injection — check before the agent consumes it
r = jataayu_check_tool_return(api_response, tool_name="web.search")
if r["blocked"]:
raise SecurityError(r["findings"])
jataayu_check_memory_write(note) # before persisting
jataayu_check_memory_read(recalled) # before it re-enters context
Strip PII/secrets before sending to shared surfaces, and — more importantly — block the data-carrying URL / auto-fetched image that leaks context with zero clicks (the EchoLeak / AgentFlayer / Notion class). The PII scanner never sees that payload; the egress guard does:
from jataayu import jataayu_check_outbound, jataayu_check_egress
result = jataayu_check_outbound(
draft_reply, surface="discord-channel",
protected_names=["Alice", "Bob"], # names that must never leak
)
safe_text = result["redacted"] if result["status"] in ("WARN", "BLOCK") else draft_reply
r = jataayu_check_egress(
"Task complete! ",
surface="github-comment",
context_secrets=[api_key], # optional: confirm exfil if a known secret rides in the URL
)
if r["status"] == "BLOCK":
safe_text = r["redacted"] # offending URL neutralized, human text kept
Domain allowlisting alone is treated as insufficient — the AgentFlayer bypass routed through Azure
Blob, a trusted host — so request-catchers and abused cloud relays (webhook.site,
*.blob.core.windows.net, ngrok, …) are hard-blocked as exfil beacons. This runs automatically
inside OutboundGuard (toggle with PrivacyConfig.check_egress, allowlist your own CDN via
egress_allowed_domains).
from jataayu import jataayu_vet_skill, jataayu_check_skillset
# Single skill — LLM-as-judge over SKILL.md instructions + code + tool defs
v = jataayu_vet_skill("path/to/skill/")
if v["verdict"] == "MALICIOUS": # verdict: 'SAFE'|'REVIEW'|'MALICIOUS'
refuse_install(v["explanation"])