
AIX360
Interpretability and explainability of data and machine learning models

Interpretability and explainability of data and machine learning models

An experimentation and research platform to investigate the interaction of automated agents in an abstract simulated network environments.

A unified framework for privacy-preserving data analysis and machine learning


An alignment auditing agent capable of quickly exploring alignment hypothesis

List of tools & datasets for anomaly detection on time-series data.

A Deep Learning Approach for Password Guessing (https://arxiv.org/abs/1709.00440)

The world's most sophisticated street level image geolocation software

NFStream: a Flexible Network Data Analysis Framework.

Re-play Security Events

Curated list of backdoor learning papers, surveys, and toolboxes, organizing poisoning-based attacks and defenses in deep learning for researchers…

Automated behavioral evaluation framework for LLMs that generates diverse test scenarios to probe for sycophancy, bias, and other safety-relevant…

Standardized adversarial robustness benchmark with a public leaderboard and downloadable model zoo for evaluating ML models against Lp attacks and…

An Open-Source Package for Textual Adversarial Attack.

The repository that contains the algorithms for generating domain names, dictionaries of malicious domain names. Developed to research the…

AI / LLM Red Team Field Manual & Consultant’s Handbook

Fully automatic censorship removal for language models

I replicated Ng's RYS method and found that duplicating 3 specific layers in Qwen2.5-32B boosts reasoning by 17% and duplicating layers 12-14 in…