
garak
Modular LLM vulnerability scanner that probes for hallucination, data leakage, prompt injection, jailbreaks, and toxicity using static, dynamic, and…

Modular LLM vulnerability scanner that probes for hallucination, data leakage, prompt injection, jailbreaks, and toxicity using static, dynamic, and…

A Python library for anomaly detection across tabular, time series, graph, text, image, and audio data. 60+ detectors, benchmark-backed ADEngine…

Set of tools to assess and improve LLM security.

Python library for adversarial machine learning security, enabling red and blue teams to run evasion, poisoning, extraction, and inference attacks…

Empire is a post-exploitation and adversary emulation framework that is used to aid Red Teams and Penetration Testers.

Open-source framework for red-teaming generative AI systems: automate attack prompts, score model responses, and audit behavior to identify security…

Curated collection of LLM jailbreak prompts and bypass techniques, documenting adversarial inputs that circumvent AI model safety guardrails.

An adversarial example library for constructing attacks, building defenses, and benchmarking both

Source code about machine learning and security.

Adversary simulation and Red teaming platform with AI

Interactive dashboards and libraries for responsible AI model debugging, covering error analysis, fairness, interpretability, counterfactuals, causal…

This repository contains detailed adversary simulation APT campaigns targeting various critical sectors. Each simulation includes custom tools, C2…

Real-time deepfake toolkit for penetration testing of identity verification and video conferencing systems. Supports face swap, image animation, and…

A comprehensive set of fairness metrics for datasets and machine learning models, explanations for these metrics, and algorithms to mitigate bias in…

A Python toolbox to create adversarial examples that fool neural networks in PyTorch, TensorFlow, and JAX

Interpretability and explainability of data and machine learning models

Algorithms for outlier, adversarial and drift detection

The AI Security Verification Standard (AISVS) focuses on providing developers, architects, and security professionals with a structured checklist to…