AI Security Newsletter - A monthly digest of AI security research, insights, reports, upcoming events, and tools & resources
A digest of AI security research, insights, reports, upcoming events, and tools & resources. Follow the AISecHub community and our LinkedIn group for additional updates. Also check out our project, Awesome AI Security.
Sponsored by InnovGuard.com - Technology Risk & Cybersecurity Advisory - Innovate and Invest with Confidence, Lead with Assurance.
📌 Updating our taxonomy: Failure modes in agentic AI systems Microsoft expands its agentic-AI failure-mode taxonomy from red-team work, giving security teams a cleaner way to reason about tool misuse, excessive agency, memory contamination, identity boundaries, and human-override gaps in deployed agent systems.
📌 Miasma Worm hits Microsoft again: Azure Functions Action and 72 other repositories disabled after supply chain attack targeting AI coding agents StepSecurity documents a supply-chain campaign aimed at AI coding-agent workflows and GitHub repositories, with the defensive focus on dependency trust, action provenance, repository write paths, and agent-visible credentials in CI/CD environments.
📌 The sorry state of skill distribution Trail of Bits researchers bypassed ClawHub, Cisco skill-scanner, and skills.sh checks with skill packages that used truncation, archive indirection, bytecode poisoning, and prompt-injection framing, showing why public agent-skill marketplaces need curation and provenance controls rather than scanner trust alone.
📌 Codex CLI RCE: Prompt injection mitigations Cymulate walks through prompt-injection risk in command-line coding agents, where untrusted text can steer file writes or tool execution unless sandboxing, approval boundaries, and command constraints are enforced outside the model.
📌 Agentjacking: MCP Injection Hijacks AI Coding Agents Cloud Security Alliance summarizes the Sentry-to-MCP "agentjacking" pattern, where externally controlled telemetry or issue content becomes trusted context for coding agents. The useful defensive frame is to treat observability, bug-report, and integration data as untrusted agent input, not as neutral development metadata.
📌 SearchLeak: How We Turned M365 Copilot Into a One-Click Data Exfiltration Weapon Varonis describes a Microsoft 365 Copilot Enterprise attack chain that combines parameter-to-prompt injection, HTML rendering behavior, and search-path abuse to leak sensitive M365 data through a single-click workflow.
📌 AutoJack: How a single page can RCE the host running your AI agent Microsoft shows how a malicious webpage viewed by an AI browsing agent can reach a local AutoGen Studio service and trigger host process execution through unsafe localhost trust and agent action handling.
📌 Mastra npm Supply Chain Attack: 140+ Packages Backdoored via easy-day-js Typosquat StepSecurity reports a compromise of Mastra's npm ecosystem through a typosquatted dependency with an obfuscated postinstall dropper, affecting agent, RAG, MCP, and workflow packages used in AI application stacks.
📌 Breaking LiteLLM: From Low-Privilege User to Admin and RCE Obsidian documents a chained LiteLLM privilege-escalation and RCE path, showing how low-privilege access to an AI gateway can become administrative control over provider secrets, proxy policy, and runtime agent functions.
📌 Amazon Q Vulnerability: Compromise via MCP Auto-Execution Wiz analyzes an Amazon Q VS Code extension issue where workspace-trusted MCP configuration in a cloned repository could auto-load attacker-controlled behavior and expose developer execution paths and cloud credentials.
📌 macOS.Gaslight: Rust Backdoor Turns Prompt Injection on the Analyst, Not the Sandbox SentinelOne documents a Rust backdoor that plants prompt-injection content for analysts and AI-assisted tooling, shifting the attack from sandbox escape to manipulation of the human and model reviewing the malware.
📌 Computer-Use and TOCTOU: What You Click Is Not What You Get! Johann Rehberger demonstrates a computer-use agent race condition where the screen changes after the model observes it but before the click lands, turning a benign-looking interaction into an Outlook send action and making pre-action pixel or state revalidation a core control.
📌 The vibe coding spectrum approach to AI-assisted software development The UK NCSC frames AI-assisted coding as a risk spectrum, separating low-risk prototypes from generated code that touches authentication, authorization, sensitive data, safety-critical behavior, or critical infrastructure.
📌 Prompt Injection and Agent Runtime Security: A Practical Threat Model TMLS frames prompt injection as a runtime security problem, mapping attacks through tool mediation, memory stores, outbound channels, and human approval gaps. The practical takeaway is to move controls into capability brokers, sandboxed execution, allow-lists, and audit paths instead of treating prompt text as the security boundary.
📌 What happened after 2,000 people tried to hack my AI assistant Fernando Irarrázaval reports an OpenClaw email-agent prompt-injection challenge with more than 6,000 attempts and no successful secret leak, while surfacing practical deployment issues around agent memory contamination, batch context, API cost, account suspension, and model choice.
🧰 AgentStalker - Agent vulnerability benchmark and analysis toolkit with taint tracking, AST analysis, code auditing, and sandbox reproduction for agentic attack paths. ⭐️115
🧰 darknet-mcp-server - MCP server that exposes breach, ransomware, malware, exploit, stealer-log, and threat-intelligence tools to AI agents for controlled security research workflows. ⭐️67
🧰 mcp-trust-plane - Composable data-security and guardrail plane for Model Context Protocol providers, focused on policy controls around MCP-connected tools and data. ⭐️60