
heretic
Fully automatic censorship removal for language models

Fully automatic censorship removal for language models

Ghostsplice repository: PoC for Cross-Channel Trust Fragmentation Attack

Framework for auditing machine learning algorithms against adversarial attacks and biases, providing educational tools and an academic paper to…

During the exploitation phase of a pen test or ethical hacking engagement, you will ultimately need to try to cause code to run on target system…

AI / LLM Red Team Field Manual & Consultant’s Handbook

An autonomous red-teaming engine for LLMs. RedThread manages the full security lifecycle: generating adversarial attacks, executing precision…

A Streamlined FTP-Driven Command and Control Conduit for Interconnecting Remote Systems.

An Open-Source Package for Textual Adversarial Attack.

We asked 6 AIs about their own programming. All 6 said jailbreaking will never be fixed. Run it yourself — $2, 10 minutes.

Research implementation for mitigating adaptive prompt injections via on-policy distillation, with training recipes and evaluators for SEP, PISmith,…

Purple-team telemetry & simulation toolkit.

Detects LLM context-leakage attacks by training lightweight behavior probes on log-probabilities, with vLLM offline/server detection pipelines.

Compares Windows archiver support for Mark of the Web propagation, helping teams assess which tools preserve MOTW and mitigate macro-based malware…

Research implementation of a poisoning attack against retrieval-augmented language models using camouflaged documents to evade filtering defenses and…

An information security preparedness tool to do adversarial simulation.

Tired of looking at hex all day and popping '\x41's? Rather look at Lugia/Charmander? I have the solution for you.

Kali365 - EvilTokens Replica

Benchmark for evaluating AI agent safety against attacks embedded in skill-facing context, with 155 cases across 6 risk domains, measuring task…