
heretic
Fully automatic censorship removal for language models

Fully automatic censorship removal for language models

Ghostsplice repository: PoC for Cross-Channel Trust Fragmentation Attack

Research implementation for mitigating adaptive prompt injections via on-policy distillation, with training recipes and evaluators for SEP, PISmith,…

Detects LLM context-leakage attacks by training lightweight behavior probes on log-probabilities, with vLLM offline/server detection pipelines.

A Python toolbox to create adversarial examples that fool neural networks in PyTorch, TensorFlow, and JAX

Open-source framework for red-teaming generative AI systems: automate attack prompts, score model responses, and audit behavior to identify security…

Protection against Model Serialization Attacks

Interactive dashboards and libraries for responsible AI model debugging, covering error analysis, fairness, interpretability, counterfactuals, causal…

Interpretability and explainability of data and machine learning models

Threadless Process Injection using remote function hooking.

Train, evaluate, and explore neural networks with built-in adversarial robustness tools, including PGD attacks, adversarial training, and robust…

C# Reflective loader for unmanaged binaries.

☁️ ⚡ Granular, Actionable Adversary Emulation for the Cloud

Real-time deepfake toolkit for penetration testing of identity verification and video conferencing systems. Supports face swap, image animation, and…

A toolset to make a system look as if it was the victim of an APT attack

AI Red Teaming playground labs to run AI Red Teaming trainings including infrastructure.

An alignment auditing agent capable of quickly exploring alignment hypothesis

Red Team K8S Adversary Emulation Based on kubectl