
awesome-ai-security
A collection of awesome resources related AI security

A collection of awesome resources related AI security

Adversary Emulation Framework

Self-Defeating Audits: reproducible lab showing a low-privilege PostgreSQL role reversibly blinding a trigger-based auditor + poisoning attribution…

Open-source AI security platform providing perimeter defense for LLMs and AI agents through swarm analysis, policy enforcement, adversarial testing,…

Detects LLM context-leakage attacks by training lightweight behavior probes on log-probabilities, with vLLM offline/server detection pipelines.

LLM security testing framework for detecting prompt injection, jailbreaks, and adversarial attacks — 190+ probes, 28 providers, single Go binary

Proof-of-concept exploit for CVE-2026-44578 that reproduces the vulnerable condition, enabling security researchers to validate affected systems and…

Proof-of-concept exploit for CVE-2026-73292: CSRF attack on Semaphore UI password change endpoint, serving a malicious page that silently resets an…

A curated list of AI Security materials and resources for Pentesters, Bug Hunters, and Security Researchers.

Spawns macOS programs through launchd's private XPC interface without execing them, making EDR record launchd as parent. Supports one-shot,…

Local white-box gradient attacks for open-weight LLMs: GCG/PEZ suffix search, layer saliency, weight snapshots, and rank-1 suffix-to-delta fitting…

Application-scoped Windows network brownouts in native C and BOF form

CVE-2026-6765 · Test only FormAutofill handlers exposed in Firefox

Fully automatic censorship removal for language models

Evades LLM text watermarks by injecting Unicode variation selectors; includes the SynthID generator, mean-g detector, normalization defenses, and…

Research-only AI watermark robustness toolkit: local reverse proxy strips C2PA/EXIF/XMP, Unicode, image/audio stego, OOXML/PDF metadata, and scans…

Strip multi-vendor AI provenance marks: Unicode text hygiene, statistical rewrite hooks, and C2PA/metadata from PNG/JPEG/SVG/PDF/DOCX/HTML/MD

PoC demonstrating quadratic DoS in Elixir html_sanitize_ex via crafted HTML; includes timing benchmarks, remote exploitation curl, and verification…