
Archived
rebuff
Multi-layered prompt injection detector for AI applications using heuristics, LLM-based analysis, vectorDB attack signatures, and canary token leak…
adversarial-attackai-securityanomaly-detection+2

Evades LLM text watermarks by injecting Unicode variation selectors; includes the SynthID generator, mean-g detector, normalization defenses, and…

Two-stage prompt-injection and jailbreak detector: regex gates plus a quantised DeBERTa-v3 ONNX classifier, with image, document, and audio support.…

StealthRL: RL framework for adversarially paraphrasing AI text to stress-test detector robustness.