
heretic
Fully automatic censorship removal for language models

Fully automatic censorship removal for language models

LLM security testing framework for detecting prompt injection, jailbreaks, and adversarial attacks — 190+ probes, 28 providers, single Go binary

Research implementation of a poisoning attack against retrieval-augmented language models using camouflaged documents to evade filtering defenses and…

An alignment auditing agent capable of quickly exploring alignment hypothesis

A diagnostic framework for measuring LLM vulnerability to Affective Contextual Erosion (ACE) and related liminal attack vectors. **Delirium** is not…

Fully automatic censorship removal for language models

AI / LLM Red Team Field Manual & Consultant’s Handbook

[ICLR 2026] - Official repo for the paper: "RedBench: A Universal Dataset for Comprehensive Red Teaming of Large Language Models"

Evaluation framework that tests whether large language models follow invisible Unicode-encoded instructions embedded in normal-looking text, with…

A novel adversarial attack on LLM based on the Exponentiated Gradient Descent technique.

Code Implementation of "Pattern Enhanced Multi-Turn Jailbreaking: Exploiting Structural Vulnerabilities in Large Language Models"

Research code and experiments for defending tool-integrated LLM agents against adversarial attacks, extending Agent Security Bench with new defense…

A high-severity prompt injection flaw in Claude AI proves that even the smartest language models can be turned into weapons — all with a few lines of…

This repository provides the official implementation of POISONCRAFT: Practical Poisoning of Retrieval-Augmented Generation for Large Language Models.

A productionized greedy coordinate gradient (GCG) attack tool for large language models (LLMs)

Syntactic Ghost: An Imperceptible General-purpose Backdoor Attacks on Pre-trained Language Models