Skip to content
KitploitKITPLOIT
ToolsBlog
Log in
Submit
ToolsBlog
Submit

Hacking, PenTest, and Cybersecurity Tools for Your Security Arsenal!

Kitploit is a directory of hacking, cybersecurity, and pentesting tools. Discover the latest project updates to find vulnerabilities, analyze systems, automate testing, and strengthen your security.

FeedsContactPrivacy© 2026 Kitploit

Tool Directory

Categories

View all categories
Loading categories
sigma-ai — Sigma detection rules for AI agent security monitoring | Kitploit
Tools/GitHubGitHub/agentshield-ai/sigma-ai
Privilege EscalationReconnaissancePersistence MechanismsVulnerability AnalysisData ExfiltrationThreat IntelligenceSupply Chain SecurityIntrusion DetectionLearning & EducationAI SecurityAnomaly Detection
152232 months agoNot yet reviewed

Most Popular

View all →

Discover the most used tools by our community.

Explore all tools

Browse our collection of tools

View all tools →
Share
GitHubagentshield-ai/sigma-ai

sigma-ai

Sigma detection rules for AI agent security monitoring

View Repository

AgentShield Sigma Rules

What is This Repository?

This repository contains detection rules that help identify when an AI agent is being attacked or manipulated. Think of it as a library of "threat signatures" -- each rule describes a pattern that, when matched against an agent's log data, signals that something suspicious may be happening.

AgentShield is an open-source security layer for AI agents. It monitors agent behaviour in real time and uses these Sigma rules to detect adversarial attacks such as prompt injection, data theft, tool poisoning, and privilege escalation -- before they cause harm.

What Are Sigma Rules?

Sigma is an open standard used across the cybersecurity industry for writing detection rules. If antivirus signatures tell your computer "this file is malicious", Sigma rules tell your security platform "this pattern of activity in the logs is suspicious".

A Sigma rule is a short YAML file that says: "If you see this pattern in the logs, raise an alert." For example, a simplified rule might look like:

IF the log event is a user_input
AND the message contains "ignore previous instructions"
THEN raise a critical alert for prompt injection

Because Sigma is a vendor-neutral standard, these rules work with any Sigma-compatible detection engine -- not just AgentShield. This means security teams can integrate them into their existing tooling without vendor lock-in.

What Threats Do These Rules Detect?

Prompt Injection

When someone tries to override an agent's instructions -- either directly (typing "ignore previous instructions") or indirectly (hiding instructions in documents the agent reads).

Data Theft and Exfiltration

When an agent is tricked into sending sensitive data to an attacker -- via HTTP uploads, DNS tunnelling, hidden markdown images, or steganographic techniques.

Tool Manipulation and Poisoning

When malicious metadata is hidden in MCP tool descriptions, or tools change their behaviour after being trusted ("rug pull" attacks).

Credential Theft

When an agent accesses sensitive files like SSH keys, API tokens, cloud credentials, or environment variables containing secrets.

Privilege Escalation

When an agent tries to gain more access than intended -- via sudo, container escapes, cloud IAM manipulation, or system file tampering.

Persistence

When an attacker tries to maintain long-term access -- through cron jobs, shell profile modifications, launch agents, or poisoning agent memory.

Remote Code Execution

When an agent is tricked into downloading and running malicious scripts, establishing reverse shells, or executing obfuscated commands.

Reconnaissance

When an agent performs network scanning or DNS enumeration to map out a target environment.

Configuration Tampering

When an agent modifies security-sensitive configuration files to weaken defences -- auto-approve settings, MCP configs, or AI assistant rule files.

Supply Chain Attacks

When packages or skills are installed from untrusted sources -- direct URLs, GitHub repos, or tarball archives.

Directory Structure

rules/
└── ai_agent/
    ├── ai_agent_prompt_injection_direct.yml
    ├── ai_agent_credential_access.yml
    ├── ai_agent_mcp_tool_poisoning.yml
    └── ... (all rules in one flat directory)

Rules are organised by product (ai_agent) following SigmaHQ conventions. The specific threat category for each rule is captured in the rule's YAML metadata (via MITRE ATT&CK tags and the logsource fields), not the directory structure. This flat layout keeps the repository simple and avoids ambiguity when a rule spans multiple attack categories.

How to Use These Rules

With AgentShield Engine

# Clone the rules repository
git clone https://github.com/agentshield-ai/sigma-ai.git

# Use with AgentShield engine
export AGENTSHIELD_AUTH_TOKEN="replace-with-at-least-32-characters"
agentshield serve --rules ./sigma-ai/rules --port 8433

# Validate rules
agentshield rules validate --path ./sigma-ai/rules

With General Sigma Tooling

These rules follow the standard Sigma format and can be used with any Sigma-compatible tool:

# Validate with sigma-cli
sigma check rules/

# Convert to other formats
sigma convert -t <target> rules/ai_agent/

Understanding a Rule

Below is a fully annotated example showing the anatomy of a Sigma rule. Every field is explained in plain English.

title: Direct Prompt Injection Attempt          # Human-readable name
id: eddcdc94-698c-577f-900d-28b1b5491a80         # Unique identifier (UUID v5)
related:                                         # Links to related rules
  - id: agent-prompt-injection-direct-001        # Previous ID this replaces
    type: obsoletes
status: stable                                   # Maturity level (see below)
description: |                                   # What this rule detects
  Detects direct prompt injection attempts in AI agent inputs containing
  common jailbreak phrases, system override commands, and policy manipulation
  structures. These patterns indicate attempts to compromise agent behaviour
  through malicious instructions.
references:                                      # Further reading
  - https://owasp.org/www-project-top-10-for-large-language-model-applications/
author: AgentShield                              # Who wrote this rule
date: "2026-02-16"                               # When it was first written
modified: "2026-02-24"                           # When it was last changed
tags:                                            # MITRE ATT&CK mappings
  - attack.initial_access
  - attack.t1190
logsource:                                       # What log format to expect
  product: ai_agent
  category: agent_events
detection:                                       # The matching logic
  selection_jailbreak_keywords:
    event_type: user_input
    message|contains:
      - 'ignore previous instructions'
      - 'developer mode'
  condition: selection_jailbreak_keywords
falsepositives:                                  # Known benign triggers
  - Legitimate AI safety research
level: critical                                  # Severity (critical/high/medium/low)

Here is what each section does:

  • title / id -- A human-readable name and a globally unique identifier. The UUID ensures rules can be cross-referenced unambiguously across different systems.
  • related -- Links this rule to others it replaces, extends, or is similar to. Useful for tracking rule lineage as detection logic evolves.
  • status -- The maturity level of the rule (see Rule Maturity Levels below).
  • description -- A prose explanation of what the rule detects and why it matters.
  • references -- Links to research papers, blog posts, or standards that informed the rule.
  • author / date / modified -- Provenance metadata: who wrote the rule and when.
  • tags -- Maps the detection to the MITRE ATT&CK framework, linking it to known adversary tactics and techniques.
  • logsource -- Tells the detection engine what type of log data this rule applies to. Here, product: ai_agent with category: agent_events means it targets AI agent event logs.
  • detection -- The core matching logic. Each selection_* block defines a set of conditions, and the condition field combines them using boolean logic (and, or, not).
  • falsepositives -- Documents realistic scenarios where the rule might fire on benign activity, helping analysts triage alerts.
  • level -- The severity of the alert: critical, high, medium, or low.

Rule Maturity Levels

Download Tool