
StealthRL
StealthRL: RL framework for adversarially paraphrasing AI text to stress-test detector robustness.

StealthRL: RL framework for adversarially paraphrasing AI text to stress-test detector robustness.

Research implementation for mitigating adaptive prompt injections via on-policy distillation, with training recipes and evaluators for SEP, PISmith,…

Trajectory-aware evolutionary search framework for red-teaming LLM agents over MCP servers, generating adversarial prompts to map vulnerability…

Code for 'Steering the CensorShip: Uncovering Representation Vectors for LLM "Thought" Control'

Curated reading list and taxonomy of attack and defense research for mobile on-device AI systems, covering adversarial, backdoor, model stealing, and…

Research demonstration of indirect prompt injection attacks to control autonomous LLM-based web agents, with tools for trigger optimization and…

Research code for Rubric-Induced Preference Drift (RIPD): evolutionary rubric search, benchmark-preserving selection, and DPO policy misalignment…

[CVPR 2025-ADVML] Official Repository for `Attacking Attention of Foundation Models Effectively Disrupts Downstream Tasks`

[ICCV 2025] Anti-Tamper Protection for Unauthorized Individual Image Generation

Voice-based detective interrogation game. Mistral Large 3 + Voxtral STT + ElevenLabs TTS. Built for the Mistral Worldwide Hackathon 2026.

Black-box attack framework that hijacks reasoning in agentic retrieval-augmented generation systems by injecting poisoned documents, with support for…

Proof-of-concept exploit for CVE-2026-73292: CSRF attack on Semaphore UI password change endpoint, serving a malicious page that silently resets an…

Educational Python PoC for a QUIC address-validation bypass that triggers handshake amplification, including vulnerable server simulation and attack…

Local white-box gradient attacks for open-weight LLMs: GCG/PEZ suffix search, layer saliency, weight snapshots, and rank-1 suffix-to-delta fitting…

A curated collection of resources for learning and researching LLM prompt injection attacks, defenses, and security.

A diagnostic framework for measuring LLM vulnerability to Affective Contextual Erosion (ACE) and related liminal attack vectors. **Delirium** is not…

Intentionally vulnerable machine learning model for hands-on security training. Explore common ML vulnerabilities, adversarial attacks, and defensive…

This project (PoC for now, and part of Shit Bucket) involves face detection, face recognition and adversarial input to protect avatars