
AI-Vulnerabilities-Playground
Hands-on AI security lab platform with 50+ scenarios across prompt injection, agentic system exploitation, model manipulation, and MCP trust boundary…

Hands-on AI security lab platform with 50+ scenarios across prompt injection, agentic system exploitation, model manipulation, and MCP trust boundary…

Code for 'Steering the CensorShip: Uncovering Representation Vectors for LLM "Thought" Control'

Research implementation for mitigating adaptive prompt injections via on-policy distillation, with training recipes and evaluators for SEP, PISmith,…

Curated reading list and taxonomy of attack and defense research for mobile on-device AI systems, covering adversarial, backdoor, model stealing, and…

A curated collection of resources for learning and researching LLM prompt injection attacks, defenses, and security.

Research demonstration of indirect prompt injection attacks to control autonomous LLM-based web agents, with tools for trigger optimization and…

Proof-of-concept exploit for CVE-2026-73292: CSRF attack on Semaphore UI password change endpoint, serving a malicious page that silently resets an…

Educational Python PoC for a QUIC address-validation bypass that triggers handshake amplification, including vulnerable server simulation and attack…

Local white-box gradient attacks for open-weight LLMs: GCG/PEZ suffix search, layer saliency, weight snapshots, and rank-1 suffix-to-delta fitting…

[CVPR 2025-ADVML] Official Repository for `Attacking Attention of Foundation Models Effectively Disrupts Downstream Tasks`

[ICCV 2025] Anti-Tamper Protection for Unauthorized Individual Image Generation

Voice-based detective interrogation game. Mistral Large 3 + Voxtral STT + ElevenLabs TTS. Built for the Mistral Worldwide Hackathon 2026.

Cross-chain bridge PoC for CVE-2026-23003: demonstrates message forging via missing origin chain ID using vulnerable Solidity contract and Python…

Intentionally vulnerable machine learning model for hands-on security training. Explore common ML vulnerabilities, adversarial attacks, and defensive…

A diagnostic framework for measuring LLM vulnerability to Affective Contextual Erosion (ACE) and related liminal attack vectors. **Delirium** is not…

This project (PoC for now, and part of Shit Bucket) involves face detection, face recognition and adversarial input to protect avatars

Code repository for CS5446 project exploring defenses against jailbreak attacks on large-language models it includes datasets, notebooks,…