
heretic
Fully automatic censorship removal for language models

Fully automatic censorship removal for language models

Benchmark harness measuring where prompt injection defenses fire in tool-using LLM agent pipelines, tracking canary tokens across exposed, persisted,…

Experiments for control-token chain-of-thought suppression and parser-leniency attacks on tool-using LLM agents

Linux process identity cloaking tool that spoofs comm, argv, cmdline, environ, exe path, and VMAs via an 11-phase prctl pipeline to impersonate…

Adversarial image perturbation tool that uses SAM segmentation and CLIP models to evade AI-based scam image classifiers for security research.

Evaluation framework that tests whether large language models follow invisible Unicode-encoded instructions embedded in normal-looking text, with…

Research code and experiments for defending tool-integrated LLM agents against adversarial attacks, extending Agent Security Bench with new defense…
