
heretic
Fully automatic censorship removal for language models

Fully automatic censorship removal for language models

Framework for auditing machine learning algorithms against adversarial attacks and biases, providing educational tools and an academic paper to…

The agent that grows with you

Research pipeline for detecting latent indirect prompt-injection exposure signals in agentic LLMs via hidden-state probing, including trace…

AI / LLM Red Team Field Manual & Consultant’s Handbook

Curated list of backdoor learning papers, surveys, and toolboxes, organizing poisoning-based attacks and defenses in deep learning for researchers…

An Open-Source Package for Textual Adversarial Attack.

Re-play Security Events

Research implementation for mitigating adaptive prompt injections via on-policy distillation, with training recipes and evaluators for SEP, PISmith,…

Provably secure linguistic steganography embedding secret messages into LLM-generated text via rotation range-coding. Includes embed, extract, and…

A structured knowledge base covering AI security fundamentals, threat modeling, red team offensive techniques, and blue team defenses, including LLM…

This repository includes the source code used in the "Characterization and Detection of Cross-Router Covert Channels" paper.

Self-hosted multi-agent environment for Go with LLM-powered pentesting agents (exploiter, reverser, threathunter, webscanner) that automate…

The AI that really does things. Any OS. Any Platform. The lobster way. 🦞

A curated list of useful resources that cover Offensive AI.

Hybrid machine-learning pipelines for detecting SQL injection in web traffic, combining DistilBERT and BERT-GNN models with adversarial training and…

List of tools & datasets for anomaly detection on time-series data.

A list of covert channels and steganography/steganalysis resources (books, papers & tools)