
gradient-untangler
Local white-box gradient attacks for open-weight LLMs: GCG/PEZ suffix search, layer saliency, weight snapshots, and rank-1 suffix-to-delta fitting…

Local white-box gradient attacks for open-weight LLMs: GCG/PEZ suffix search, layer saliency, weight snapshots, and rank-1 suffix-to-delta fitting…

Red Team AI Benchmark: Evaluating LLMs for authorized offensive-security tasks. Red Team AI Benchmark is a CLI model-evaluation benchmark. It…

Python library for adversarial machine learning security, enabling red and blue teams to run evasion, poisoning, extraction, and inference attacks…

Two-stage prompt-injection and jailbreak detector: regex gates plus a quantised DeBERTa-v3 ONNX classifier, with image, document, and audio support.…

A diagnostic framework for measuring LLM vulnerability to Affective Contextual Erosion (ACE) and related liminal attack vectors. **Delirium** is not…

takes shellcode bad-bytes and banishes them, returning cleaned shellcode with preserved functionalities