
OpenAttack
An Open-Source Package for Textual Adversarial Attack.

An Open-Source Package for Textual Adversarial Attack.

Open-source framework for red-teaming generative AI systems: automate attack prompts, score model responses, and audit behavior to identify security…

Python library for adversarial machine learning security, enabling red and blue teams to run evasion, poisoning, extraction, and inference attacks…

A Python toolbox to create adversarial examples that fool neural networks in PyTorch, TensorFlow, and JAX

Website defacement attack detection with deep learning

A Python library for Secure and Explainable Machine Learning Documentation available @ https://secml.gitlab.io Follow us on Twitter @…

A Python library for anomaly detection across tabular, time series, graph, text, image, and audio data. 60+ detectors, benchmark-backed ADEngine…

A novel adversarial attack on LLM based on the Exponentiated Gradient Descent technique.

Fully automatic censorship removal for language models

A comprehensive set of fairness metrics for datasets and machine learning models, explanations for these metrics, and algorithms to mitigate bias in…

Algorithms for outlier, adversarial and drift detection

Fully automatic censorship removal for language models

PISmith: Reinforcement Learning-based Red Teaming for Prompt Injection Defenses

Reverse Shell Detection with Machine Learning

Local white-box gradient attacks for open-weight LLMs: GCG/PEZ suffix search, layer saliency, weight snapshots, and rank-1 suffix-to-delta fitting…

Interactive dashboards and libraries for responsible AI model debugging, covering error analysis, fairness, interpretability, counterfactuals, causal…

image scaling attacks for multi-modal prompt injection

A Binary Genetic Traits Lexer Framework