
T-Backdoor
Research code implementing T-Backdoor, temporal-trigger backdoor attacks on spiking neural networks using rate, latency, and jitter triggers without…

Research code implementing T-Backdoor, temporal-trigger backdoor attacks on spiking neural networks using rate, latency, and jitter triggers without…
Defense framework that keeps audio-language models frozen and adds a mid-layer risk gate with late-layer safety adapters to block audio jailbreaks at…

Benchmark suite and code for detecting AI alignment failures, with 44 benchmarks across ten failure types and a zero-shot RLCD detector evaluated on…

0-day malware detection for binaries, source & scripts (that doesn't suck)

DNS Proxy that is simple and fast with not so simple features. Focused on routed DNS forwarding, filtering and parental control.

Black-box input-stage purification defense that neutralizes backdoor attacks on object detectors via corruption, diffusion reconstruction, and DBSCAN…

Benchmarking prompt injection detections for web agents.

Defensive framework that maintains a safety-focused shadow memory to detect and block prompt-injection and long-horizon threats against LLM agents…

Measuring Open Privilege in Agent Defenses

Experiments for control-token chain-of-thought suppression and parser-leniency attacks on tool-using LLM agents

Proof-of-concept that poisons MLflow registered models via the REST API, embedding a malicious pickle to trigger RCE when the model is loaded.

Proof-of-concept and research repository for CVE-2024-37054, an unsafe deserialization flaw in MLflow PyFunc model loading that can lead to remote…

Research code for red-teaming AI auto-mode monitors, including simulation evals, fuzzing, and monitor implementations for Claude Code and Codex…

Multi-format malware analysis platform combining a stealth Ring-3 Windows sandbox, static PE/PDF analyzers, ransomware key recovery, and an AI…

Adaptive two-stage Layer 4 DDoS mitigation gateway using behavioral traffic analysis, Random Forest classification, and kernel-level ipset/iptables…

Bidirectional token-classification model for PII detection and masking in text, with CLI for redaction, evaluation, and finetuning on-premises.

Research code for a gray-box trojan attack that flips a single KV-cache bit in fine-tuned LLM classifiers and measures per-class attack success rate.

Pre-execution action-auditing defense that detects and masks indirect prompt injection in tool-using LLM agents using embedding retrieval and…