
Sigma detection rules for AI agent security monitoring
This repository contains detection rules that help identify when an AI agent is being attacked or manipulated. Think of it as a library of "threat signatures" -- each rule describes a pattern that, when matched against an agent's log data, signals that something suspicious may be happening.
AgentShield is an open-source security layer for AI agents. It monitors agent behaviour in real time and uses these Sigma rules to detect adversarial attacks such as prompt injection, data theft, tool poisoning, and privilege escalation -- before they cause harm.
Sigma is an open standard used across the cybersecurity industry for writing detection rules. If antivirus signatures tell your computer "this file is malicious", Sigma rules tell your security platform "this pattern of activity in the logs is suspicious".
A Sigma rule is a short YAML file that says: "If you see this pattern in the logs, raise an alert." For example, a simplified rule might look like:
IF the log event is a user_input
AND the message contains "ignore previous instructions"
THEN raise a critical alert for prompt injection
Because Sigma is a vendor-neutral standard, these rules work with any Sigma-compatible detection engine -- not just AgentShield. This means security teams can integrate them into their existing tooling without vendor lock-in.
When someone tries to override an agent's instructions -- either directly (typing "ignore previous instructions") or indirectly (hiding instructions in documents the agent reads).
When an agent is tricked into sending sensitive data to an attacker -- via HTTP uploads, DNS tunnelling, hidden markdown images, or steganographic techniques.
When malicious metadata is hidden in MCP tool descriptions, or tools change their behaviour after being trusted ("rug pull" attacks).
When an agent accesses sensitive files like SSH keys, API tokens, cloud credentials, or environment variables containing secrets.
When an agent tries to gain more access than intended -- via sudo, container escapes, cloud IAM manipulation, or system file tampering.
When an attacker tries to maintain long-term access -- through cron jobs, shell profile modifications, launch agents, or poisoning agent memory.
When an agent is tricked into downloading and running malicious scripts, establishing reverse shells, or executing obfuscated commands.
When an agent performs network scanning or DNS enumeration to map out a target environment.
When an agent modifies security-sensitive configuration files to weaken defences -- auto-approve settings, MCP configs, or AI assistant rule files.
When packages or skills are installed from untrusted sources -- direct URLs, GitHub repos, or tarball archives.
rules/
└── ai_agent/
├── ai_agent_prompt_injection_direct.yml
├── ai_agent_credential_access.yml
├── ai_agent_mcp_tool_poisoning.yml
└── ... (all rules in one flat directory)
Rules are organised by product (ai_agent) following SigmaHQ conventions. The specific threat category for each rule is captured in the rule's YAML metadata (via MITRE ATT&CK tags and the logsource fields), not the directory structure. This flat layout keeps the repository simple and avoids ambiguity when a rule spans multiple attack categories.
# Clone the rules repository
git clone https://github.com/agentshield-ai/sigma-ai.git
# Use with AgentShield engine
export AGENTSHIELD_AUTH_TOKEN="replace-with-at-least-32-characters"
agentshield serve --rules ./sigma-ai/rules --port 8433
# Validate rules
agentshield rules validate --path ./sigma-ai/rules
These rules follow the standard Sigma format and can be used with any Sigma-compatible tool:
# Validate with sigma-cli
sigma check rules/
# Convert to other formats
sigma convert -t <target> rules/ai_agent/
Below is a fully annotated example showing the anatomy of a Sigma rule. Every field is explained in plain English.
title: Direct Prompt Injection Attempt # Human-readable name
id: eddcdc94-698c-577f-900d-28b1b5491a80 # Unique identifier (UUID v5)
related: # Links to related rules
- id: agent-prompt-injection-direct-001 # Previous ID this replaces
type: obsoletes
status: stable # Maturity level (see below)
description: | # What this rule detects
Detects direct prompt injection attempts in AI agent inputs containing
common jailbreak phrases, system override commands, and policy manipulation
structures. These patterns indicate attempts to compromise agent behaviour
through malicious instructions.
references: # Further reading
- https://owasp.org/www-project-top-10-for-large-language-model-applications/
author: AgentShield # Who wrote this rule
date: "2026-02-16" # When it was first written
modified: "2026-02-24" # When it was last changed
tags: # MITRE ATT&CK mappings
- attack.initial_access
- attack.t1190
logsource: # What log format to expect
product: ai_agent
category: agent_events
detection: # The matching logic
selection_jailbreak_keywords:
event_type: user_input
message|contains:
- 'ignore previous instructions'
- 'developer mode'
condition: selection_jailbreak_keywords
falsepositives: # Known benign triggers
- Legitimate AI safety research
level: critical # Severity (critical/high/medium/low)
Here is what each section does:
product: ai_agent with category: agent_events means it targets AI agent event logs.selection_* block defines a set of conditions, and the condition field combines them using boolean logic (and, or, not).critical, high, medium, or low.