
T-MAP
Trajectory-aware evolutionary search framework for red-teaming LLM agents over MCP servers, generating adversarial prompts to map vulnerability…

Trajectory-aware evolutionary search framework for red-teaming LLM agents over MCP servers, generating adversarial prompts to map vulnerability…

Reference implementation of a multi-bit LLM watermarking scheme using coded payload spreading, unbiased reweighting, and soft-decision ECC decoding…

Research code for extracting and training safety-awareness directions in multimodal LLMs to improve refusal behavior while limiting benign-task drift.

Detects unknown jailbreak attacks in large vision-language models using hidden state analysis and autoencoders, with training and evaluation…

Research code for detecting and detoxifying backdoors in text-to-image diffusion models, with pipelines for Stable Diffusion v1.4, v1.5, XL, and 3…

Linux system-call monitor using ptrace to trace file, process, network, and memory activity, with namespace isolation and machine learning…

A list of covert channels and steganography/steganalysis resources (books, papers & tools)

Python library for adversarial machine learning security, enabling red and blue teams to run evasion, poisoning, extraction, and inference attacks…

An adversarial example library for constructing attacks, building defenses, and benchmarking both

A Python toolbox to create adversarial examples that fool neural networks in PyTorch, TensorFlow, and JAX

Ready-to-use OCR with 80+ supported languages and all popular writing scripts including Latin, Chinese, Arabic, Devanagari, Cyrillic and etc.

A curated list of useful resources that cover Offensive AI.

Open-source framework for red-teaming generative AI systems: automate attack prompts, score model responses, and audit behavior to identify security…

Interactive dashboards and libraries for responsible AI model debugging, covering error analysis, fairness, interpretability, counterfactuals, causal…

Programmable guardrails for LLM chat apps: enforce input/output rails, block jailbreaks and prompt injections, detect hallucination, and mask…

A comprehensive set of fairness metrics for datasets and machine learning models, explanations for these metrics, and algorithms to mitigate bias in…

Android Antivirus which doesn't require root, adb, ca install and cloud with many features and ways to detect more zero-day malware

Analyzes LLM internal states and 100+ attention/probability features to train classifiers that detect document poisoning attacks in RAG systems.