
LeakGauge
Detects LLM context-leakage attacks by training lightweight behavior probes on log-probabilities, with vLLM offline/server detection pipelines.

Detects LLM context-leakage attacks by training lightweight behavior probes on log-probabilities, with vLLM offline/server detection pipelines.

Self-Defeating Audits: reproducible lab showing a low-privilege PostgreSQL role reversibly blinding a trigger-based auditor + poisoning attribution…

A curated list of AI Security materials and resources for Pentesters, Bug Hunters, and Security Researchers.

A curated collection of resources for learning and researching LLM prompt injection attacks, defenses, and security.

HookChain: A new perspective for Bypassing EDR Solutions

[ICLR 2026] - Official repo for the paper: "RedBench: A Universal Dataset for Comprehensive Red Teaming of Large Language Models"


Two-stage prompt-injection and jailbreak detector: regex gates plus a quantised DeBERTa-v3 ONNX classifier, with image, document, and audio support.…

Fully automatic censorship removal for language models

Reverse Shell Detection with Machine Learning

Practical black-box adversarial packet generation against encrypted traffic classification with minimal overhead and full packet recoverability.

CEREBRO-RED v2: Advanced LLM Red Team Research Platform with PAIR Algorithm and LLM-as-a-Judge Evaluation

reverse engineering Gemini's SynthID detection

Open-source cross-modal and multimodal prompt injection test suite. 250,000+ attack payloads across text, image, document, and audio modalities.…

A high-severity prompt injection flaw in Claude AI proves that even the smartest language models can be turned into weapons — all with a few lines of…

Evaluation framework that tests whether large language models follow invisible Unicode-encoded instructions embedded in normal-looking text, with…

Fully automatic censorship removal for language models
