
OMLASP
Framework for auditing machine learning algorithms against adversarial attacks and biases, providing educational tools and an academic paper to…

Framework for auditing machine learning algorithms against adversarial attacks and biases, providing educational tools and an academic paper to…

Python library for adversarial machine learning security, enabling red and blue teams to run evasion, poisoning, extraction, and inference attacks…

Fully automatic censorship removal for language models

A Python pickling decompiler and static analyzer

Research pipeline for detecting latent indirect prompt-injection exposure signals in agentic LLMs via hidden-state probing, including trace…

The agent that grows with you

AI / LLM Red Team Field Manual & Consultant’s Handbook

An Open-Source Package for Textual Adversarial Attack.

Research implementation for mitigating adaptive prompt injections via on-policy distillation, with training recipes and evaluators for SEP, PISmith,…

Curated list of backdoor learning papers, surveys, and toolboxes, organizing poisoning-based attacks and defenses in deep learning for researchers…

Collection of CVE(work) on tenserflow binary pwning it

🐢 Open-Source Evaluation & Testing library for LLM Agents

Research code for detecting and detoxifying backdoors in text-to-image diffusion models, with pipelines for Stable Diffusion v1.4, v1.5, XL, and 3…

Programmable guardrails for LLM chat apps: enforce input/output rails, block jailbreaks and prompt injections, detect hallucination, and mask…

Detects LLM context-leakage attacks by training lightweight behavior probes on log-probabilities, with vLLM offline/server detection pipelines.

Detects unknown jailbreak attacks in large vision-language models using hidden state analysis and autoencoders, with training and evaluation…

NVR with realtime local object detection for IP cameras

Algorithms for outlier, adversarial and drift detection