
LeakGauge
Detects LLM context-leakage attacks by training lightweight behavior probes on log-probabilities, with vLLM offline/server detection pipelines.

Detects LLM context-leakage attacks by training lightweight behavior probes on log-probabilities, with vLLM offline/server detection pipelines.


Fully automatic censorship removal for language models

reverse engineering Gemini's SynthID detection

Automated behavioral evaluation framework for LLMs that generates diverse test scenarios to probe for sycophancy, bias, and other safety-relevant…

An alignment auditing agent capable of quickly exploring alignment hypothesis

Metasploit for machine learning.

Fully automatic censorship removal for language models

Open-source cross-modal and multimodal prompt injection test suite. 250,000+ attack payloads across text, image, document, and audio modalities.…

Evaluation framework that tests whether large language models follow invisible Unicode-encoded instructions embedded in normal-looking text, with…

Reverse Shell Detection with Machine Learning

Practical black-box adversarial packet generation against encrypted traffic classification with minimal overhead and full packet recoverability.

Two-stage prompt-injection and jailbreak detector: regex gates plus a quantised DeBERTa-v3 ONNX classifier, with image, document, and audio support.…

Production AI defense with 7-layer protection: mathematical constraints, object-capability access, distributed O2 consensus, SVETILO ethics. First…