
ai-llm-red-team-handbook
AI / LLM Red Team Field Manual & Consultant’s Handbook

AI / LLM Red Team Field Manual & Consultant’s Handbook

Bypass restricted and censored content on AI chat prompts 😈

PoC code from DEF CON 25 presentation

Anti-LLM obfuscation via finger counting

Malware Mutation Using Reinforcement Learning and Generative Adversarial Networks

reverse engineering SynthID for text

PISmith: Reinforcement Learning-based Red Teaming for Prompt Injection Defenses

A Python library for Secure and Explainable Machine Learning Documentation available @ https://secml.gitlab.io Follow us on Twitter @…

Open-source AI security platform providing perimeter defense for LLMs and AI agents through swarm analysis, policy enforcement, adversarial testing,…

A high-severity prompt injection flaw in Claude AI proves that even the smartest language models can be turned into weapons — all with a few lines of…

CEREBRO-RED v2: Advanced LLM Red Team Research Platform with PAIR Algorithm and LLM-as-a-Judge Evaluation

Hands-on AI security lab platform with 50+ scenarios across prompt injection, agentic system exploitation, model manipulation, and MCP trust boundary…

Meet Eclipse the only jailbreak that moonwalks around ChatGPT 4o.

Automated framework for hijacking safety reasoning in large reasoning models via simulated reasoning traces and iterative prompt refinement to…

This repository provides the official implementation of POISONCRAFT: Practical Poisoning of Retrieval-Augmented Generation for Large Language Models.

A prompt injection in a code‑review bot that executes AI‑generated fixes in a sandbox. The sandbox uses a blacklist to prevent dangerous commands,…

The code for ACM MM2024 (Multimodal Unlearnable Examples: Protecting Data against Multimodal Contrastive Learning)

Research demonstration of indirect prompt injection attacks to control autonomous LLM-based web agents, with tools for trigger optimization and…