
AGAS
LLM-driven agentic group shilling attack framework that manipulates black-box collaborative-filtering recommender rankings using adaptive multi-role…

LLM-driven agentic group shilling attack framework that manipulates black-box collaborative-filtering recommender rankings using adaptive multi-role…

Browser-hooking framework for authorized red teams and educators. Hooks browsers via XSS, provides interactive post-exploitation control, blind-XSS…

Enterprise AI agent security toolkit providing pre-flight auditing, configuration hardening, runtime threat detection, and active defense against…

Free security-baseline rule for Claude Code, Codex, and Cursor: treats MCP tool descriptions as untrusted input (OWASP MCP Top 10 MCP03,…

Reproduces CVE-2026-44246, a prompt injection vulnerability in nnU-Net's GitHub Actions triage agent, demonstrating how issue content is inlined into…

Autonomous AI Driven OSINT & security-research desktop agent (macOS/Windows) that builds a live knowledge graph. Authorized use only

A deterministic harness and handbook for autonomous offensive LLM agents, enforcing authorization, scope, and evidence gates to ensure reproducible…

A structured knowledge base covering AI security fundamentals, threat modeling, red team offensive techniques, and blue team defenses, including LLM…

Benchmark for evaluating safety risks of computer-using agents, with 104 realistic misuse scenarios across seven malicious categories, supporting…

A contextual security auditing system for research artifacts

Reference implementation of a multi-bit LLM watermarking scheme using coded payload spreading, unbiased reweighting, and soft-decision ECC decoding…

Research code for extracting and training safety-awareness directions in multimodal LLMs to improve refusal behavior while limiting benign-task drift.

Defense framework enforcing routed origin policy to prevent indirect prompt injection in tool-using LLM agents, with deterministic origin checks and…

Research implementation of a poisoning attack against retrieval-augmented language models using camouflaged documents to evade filtering defenses and…

Benchmark for evaluating AI agent safety against attacks embedded in skill-facing context, with 155 cases across 6 risk domains, measuring task…

Technical disclosure: Credential fabrication via XML tag injection in Claude Sonnet 4.6. Reported June 14 2026, patched June 20 2026.

Semantic Watermarking with Order-Robust Detection over Sub-sentence Units

AI Infrastructure Vulnerability Research. CVE-2026-78906: Prompt injection and memory exfiltration in OpenAI's ChatGPT API.