
responsible-ai-toolbox
Interactive dashboards and libraries for responsible AI model debugging, covering error analysis, fairness, interpretability, counterfactuals, causal…

Interactive dashboards and libraries for responsible AI model debugging, covering error analysis, fairness, interpretability, counterfactuals, causal…

Interpretability and explainability of data and machine learning models

A comprehensive set of fairness metrics for datasets and machine learning models, explanations for these metrics, and algorithms to mitigate bias in…

A unified framework for privacy-preserving data analysis and machine learning

An alignment auditing agent capable of quickly exploring alignment hypothesis


The world's most sophisticated street level image geolocation software

A Deep Learning Approach for Password Guessing (https://arxiv.org/abs/1709.00440)

Re-play Security Events

Automated behavioral evaluation framework for LLMs that generates diverse test scenarios to probe for sycophancy, bias, and other safety-relevant…

An Open-Source Package for Textual Adversarial Attack.

The ultimate steganography and digital forensics toolkit. Hide and extract data across images, audio, video, documents, and network packets, or run…

AI / LLM Red Team Field Manual & Consultant’s Handbook

LSTM-based classifier for detecting domain generation algorithm (DGA) domains, with Keras implementations of neural network and bigram models for DNS…

AutoPentest-DRL: Automated Penetration Testing Using Deep Reinforcement Learning

Fully automatic censorship removal for language models

A deep learning toolkit for log-based anomaly detection

The repository that contains the algorithms for generating domain names, dictionaries of malicious domain names. Developed to research the…