
bloom
Automated behavioral evaluation framework for LLMs that generates diverse test scenarios to probe for sycophancy, bias, and other safety-relevant…

Automated behavioral evaluation framework for LLMs that generates diverse test scenarios to probe for sycophancy, bias, and other safety-relevant…

An Open-Source Package for Textual Adversarial Attack.

Bluetooth keystroke injection exploit PoCs for CVE-2023-45866, CVE-2024-21306, and CVE-2024-0230 targeting Android, Linux, macOS, and iOS via…

Fully automatic censorship removal for language models

Standardized adversarial robustness benchmark with a public leaderboard and downloadable model zoo for evaluating ML models against Lp attacks and…

Hands-on DEFCON workshop materials for killing and silencing EDR agents: lab setup, BYOVD, custom C/C++ evasion tooling, and reverse engineering.

Compares Windows archiver support for Mark of the Web propagation, helping teams assess which tools preserve MOTW and mitigate macro-based malware…

A font-based deception tool for red teaming, security research, and whatever else.

LLM security testing framework for detecting prompt injection, jailbreaks, and adversarial attacks — 190+ probes, 28 providers, single Go binary

Tools and PoCs for Windows syscall investigation.

Weaponizing to get NT SYSTEM for Privileged Directory Creation Bugs with Windows Error Reporting

Alignment-research scaffold (autoresearch-style) for LLM guardrails: search over a single policy.md surface

A curated list of AI Security materials and resources for Pentesters, Bug Hunters, and Security Researchers.

A Streamlined FTP-Driven Command and Control Conduit for Interconnecting Remote Systems.

[ICLR 2026] - Official repo for the paper: "RedBench: A Universal Dataset for Comprehensive Red Teaming of Large Language Models"

A Python library for Secure and Explainable Machine Learning Documentation available @ https://secml.gitlab.io Follow us on Twitter @…

StealthRL: RL framework for adversarially paraphrasing AI text to stress-test detector robustness.

Framework for auditing machine learning algorithms against adversarial attacks and biases, providing educational tools and an academic paper to…