
Learning-to-Detect
Detects unknown jailbreak attacks in large vision-language models using hidden state analysis and autoencoders, with training and evaluation…

Detects unknown jailbreak attacks in large vision-language models using hidden state analysis and autoencoders, with training and evaluation…

Research code for detecting and detoxifying backdoors in text-to-image diffusion models, with pipelines for Stable Diffusion v1.4, v1.5, XL, and 3…

Research pipeline for detecting latent indirect prompt-injection exposure signals in agentic LLMs via hidden-state probing, including trace…

Research implementation for mitigating adaptive prompt injections via on-policy distillation, with training recipes and evaluators for SEP, PISmith,…

ML-driven threat detection and continuous monitoring platform built for federal zero trust architectures.

A Python pickling decompiler and static analyzer

Linux system-call monitor using ptrace to trace file, process, network, and memory activity, with namespace isolation and machine learning…

An Evaluation Agent for Detecting Misinformation and Knowledge Poisoning in Retrieval-Augmented Generation Systems.

A list of covert channels and steganography/steganalysis resources (books, papers & tools)

Python library for adversarial machine learning security, enabling red and blue teams to run evasion, poisoning, extraction, and inference attacks…

An adversarial example library for constructing attacks, building defenses, and benchmarking both

A Python toolbox to create adversarial examples that fool neural networks in PyTorch, TensorFlow, and JAX

Protection against Model Serialization Attacks

Algorithms for outlier, adversarial and drift detection

Ready-to-use OCR with 80+ supported languages and all popular writing scripts including Latin, Chinese, Arabic, Devanagari, Cyrillic and etc.

Curated list of backdoor learning papers, surveys, and toolboxes, organizing poisoning-based attacks and defenses in deep learning for researchers…

List of tools & datasets for anomaly detection on time-series data.

Train, evaluate, and explore neural networks with built-in adversarial robustness tools, including PGD attacks, adversarial training, and robust…