
better_opts_attacks
Attack framework for breaking fine-tuning based prompt injection defenses (SecAlign, SecAlign++, StruQ) using architecture-aware adversarial attacks…

Attack framework for breaking fine-tuning based prompt injection defenses (SecAlign, SecAlign++, StruQ) using architecture-aware adversarial attacks…

Playbook-based adversary simulation framework that compiles JSON-defined attack paths into position-independent shellcode payloads for validating…

Open-source framework for red-teaming generative AI systems: automate attack prompts, score model responses, and audit behavior to identify security…

A diagnostic framework for measuring LLM vulnerability to Affective Contextual Erosion (ACE) and related liminal attack vectors. **Delirium** is not…

Purple Team Exercise Framework

Automated behavioral evaluation framework for LLMs that generates diverse test scenarios to probe for sycophancy, bias, and other safety-relevant…

LLM security testing framework for detecting prompt injection, jailbreaks, and adversarial attacks — 190+ probes, 28 providers, single Go binary

StealthRL: RL framework for adversarially paraphrasing AI text to stress-test detector robustness.

Framework for auditing machine learning algorithms against adversarial attacks and biases, providing educational tools and an academic paper to…

Implementation of paper "DeeCLIP: A Robust and Generalizable Transformer-Based Framework for Detecting AI-Generated Images"

Automated framework for hijacking safety reasoning in large reasoning models via simulated reasoning traces and iterative prompt refinement to…

Evaluation framework that tests whether large language models follow invisible Unicode-encoded instructions embedded in normal-looking text, with…

Benchmarking framework for evaluating computer-use AI agents against multi-step indirect prompt injection, with automatic adversarial goal…

The Python Risk Identification Tool for generative AI (PyRIT) is an open source framework built to empower security professionals and engineers to…

Metasploit for machine learning.