
reverse-captcha-eval
Evaluation framework that tests whether large language models follow invisible Unicode-encoded instructions embedded in normal-looking text, with…
adversarial-attackai-securityctf+3
11

Evaluation framework that tests whether large language models follow invisible Unicode-encoded instructions embedded in normal-looking text, with…

Modular LLM vulnerability scanner that probes for hallucination, data leakage, prompt injection, jailbreaks, and toxicity using static, dynamic, and…