
heretic
Fully automatic censorship removal for language models

Fully automatic censorship removal for language models

Ghostsplice repository: PoC for Cross-Channel Trust Fragmentation Attack

Framework for auditing machine learning algorithms against adversarial attacks and biases, providing educational tools and an academic paper to…

We asked 6 AIs about their own programming. All 6 said jailbreaking will never be fixed. Run it yourself — $2, 10 minutes.

Kali365 - EvilTokens Replica

During the exploitation phase of a pen test or ethical hacking engagement, you will ultimately need to try to cause code to run on target system…

A toolset to make a system look as if it was the victim of an APT attack

AI / LLM Red Team Field Manual & Consultant’s Handbook

A Streamlined FTP-Driven Command and Control Conduit for Interconnecting Remote Systems.

An autonomous red-teaming engine for LLMs. RedThread manages the full security lifecycle: generating adversarial attacks, executing precision…

An Open-Source Package for Textual Adversarial Attack.

Research implementation for mitigating adaptive prompt injections via on-policy distillation, with training recipes and evaluators for SEP, PISmith,…

Purple-team telemetry & simulation toolkit.

Detects LLM context-leakage attacks by training lightweight behavior probes on log-probabilities, with vLLM offline/server detection pipelines.

Research implementation of a poisoning attack against retrieval-augmented language models using camouflaged documents to evade filtering defenses and…

Compares Windows archiver support for Mark of the Web propagation, helping teams assess which tools preserve MOTW and mitigate macro-based malware…

Open-source framework for red-teaming generative AI systems: automate attack prompts, score model responses, and audit behavior to identify security…

An information security preparedness tool to do adversarial simulation.