
bordair-detector
Two-stage prompt-injection and jailbreak detector: regex gates plus a quantised DeBERTa-v3 ONNX classifier, with image, document, and audio support.…

Two-stage prompt-injection and jailbreak detector: regex gates plus a quantised DeBERTa-v3 ONNX classifier, with image, document, and audio support.…

Demonstrates how a malicious Python package executes arbitrary commands during pip install via setup.py, highlighting PyPI supply chain and…

Splunk SIEM lab simulating and detecting CVE-2021-34527 (PrintNightmare) exploitation using Sysmon, Windows Event logs, and custom SPL detection…

A curated collection of resources for learning and researching LLM prompt injection attacks, defenses, and security.

Multi-stage prompt injection technique that bypasses LLM safety alignment via identity reassignment, refusal suppression, and output coercion,…

A diagnostic framework for measuring LLM vulnerability to Affective Contextual Erosion (ACE) and related liminal attack vectors. **Delirium** is not…

Intentionally vulnerable machine learning model for hands-on security training. Explore common ML vulnerabilities, adversarial attacks, and defensive…

This project (PoC for now, and part of Shit Bucket) involves face detection, face recognition and adversarial input to protect avatars

Cross-chain bridge PoC for CVE-2026-23003: demonstrates message forging via missing origin chain ID using vulnerable Solidity contract and Python…

Code repository for CS5446 project exploring defenses against jailbreak attacks on large-language models it includes datasets, notebooks,…

Research code for Rubric-Induced Preference Drift (RIPD): evolutionary rubric search, benchmark-preserving selection, and DPO policy misalignment…

Multi-layered prompt injection detector for AI applications using heuristics, LLM-based analysis, vectorDB attack signatures, and canary token leak…

Research code and experiments for defending tool-integrated LLM agents against adversarial attacks, extending Agent Security Bench with new defense…

This repository contains the official implementation of the paper "[Safety in Batches? Understanding and Mitigating Safety Failures in Batch…

Research code for a gray-box trojan attack that flips a single KV-cache bit in fine-tuned LLM classifiers and measures per-class attack success rate.

The Python Risk Identification Tool for generative AI (PyRIT) is an open source framework built to empower security professionals and engineers to…

A privacy-first app that strips AI watermarks from content you own.

Automated Adversary Emulation Platform