
heretic
Fully automatic censorship removal for language models

Fully automatic censorship removal for language models

Bypass restricted and censored content on AI chat prompts 😈

Ghostsplice repository: PoC for Cross-Channel Trust Fragmentation Attack

Framework for auditing machine learning algorithms against adversarial attacks and biases, providing educational tools and an academic paper to…

Open-source framework for red-teaming generative AI systems: automate attack prompts, score model responses, and audit behavior to identify security…

We asked 6 AIs about their own programming. All 6 said jailbreaking will never be fixed. Run it yourself — $2, 10 minutes.

The TTPForge is a Cybersecurity Framework for developing, automating, and executing attacker Tactics, Techniques, and Procedures (TTPs).

Kali365 - EvilTokens Replica

The exploit server for out-of-band findings. Point a target at a domain you own. Every HTTP request and every email it sends back lands in a…

During the exploitation phase of a pen test or ethical hacking engagement, you will ultimately need to try to cause code to run on target system…

A toolset to make a system look as if it was the victim of an APT attack

An alignment auditing agent capable of quickly exploring alignment hypothesis

Causal context attribution and rule-based monitor LLM defense against indirect prompt injection in LLM agents, achieving state-of-art performance on…

Red Team K8S Adversary Emulation Based on kubectl

A Streamlined FTP-Driven Command and Control Conduit for Interconnecting Remote Systems.

Scripted framework for simulating over 50 MITRE ATT&CK techniques to test blue team detection capabilities. Includes Python scripts and a compiled…

An autonomous red-teaming engine for LLMs. RedThread manages the full security lifecycle: generating adversarial attacks, executing precision…

An Open-Source Package for Textual Adversarial Attack.