
LeakGauge
Detects LLM context-leakage attacks by training lightweight behavior probes on log-probabilities, with vLLM offline/server detection pipelines.

Detects LLM context-leakage attacks by training lightweight behavior probes on log-probabilities, with vLLM offline/server detection pipelines.

Fully automatic censorship removal for language models

An alignment auditing agent capable of quickly exploring alignment hypothesis

Production AI defense with 7-layer protection: mathematical constraints, object-capability access, distributed O2 consensus, SVETILO ethics. First…

Open-source cross-modal and multimodal prompt injection test suite. 250,000+ attack payloads across text, image, document, and audio modalities.…

Two-stage prompt-injection and jailbreak detector: regex gates plus a quantised DeBERTa-v3 ONNX classifier, with image, document, and audio support.…

reverse engineering Gemini's SynthID detection

Fully automatic censorship removal for language models

Practical black-box adversarial packet generation against encrypted traffic classification with minimal overhead and full packet recoverability.

Reverse Shell Detection with Machine Learning

Automated behavioral evaluation framework for LLMs that generates diverse test scenarios to probe for sycophancy, bias, and other safety-relevant…

Evaluation framework that tests whether large language models follow invisible Unicode-encoded instructions embedded in normal-looking text, with…


Metasploit for machine learning.