
SkillPoison
Research pipeline that constructs verified successful experiences and organizes them into formation records to progressively poison model skills via…

Research pipeline that constructs verified successful experiences and organizes them into formation records to progressively poison model skills via…

Research implementation of Hop-Decayed Influence (HDI) and the 3S attack framework, exposing structural auxiliary indexing vulnerabilities in…

Proof-of-concept demos and research on indirect prompt injection attacks against application-integrated LLMs, covering data exfiltration, remote…

Refusal localizes, the damage relocates, safety layers under few-sample fine-tuning

Curated index of deep learning latency, energy-latency & timing (availability) attacks and defenses — companion to the ACM Computing Surveys…

Curated collection of LLM jailbreak prompts and bypass techniques, documenting adversarial inputs that circumvent AI model safety guardrails.

Automated adversary emulation (Caldera) against an AD lab to validate Sigma detection coverage and map results to MITRE ATT&CK.

Tracker of publicly reported prompt-injection techniques, broken down by delivery method, encoding, and propagation behavior, with confirmed models,…

Experiments for control-token chain-of-thought suppression and parser-leniency attacks on tool-using LLM agents

Research code reproducing multi-turn LLM jailbreak experiments (FITD, MRCJ, ActorAttack, X-Teaming) from the SoK intent-oriented systematization…

This repository contains the official implementation of the paper "[Safety in Batches? Understanding and Mitigating Safety Failures in Batch…

Black-box attack framework that hijacks reasoning in agentic retrieval-augmented generation systems by injecting poisoned documents, with support for…

Research implementation of a poisoning attack against retrieval-augmented language models using camouflaged documents to evade filtering defenses and…

Crystal port of GodPotato to abuse SeImpersonatePrivilege with indirect syscalls, dynamic API resolution and compile-time string obfuscation. Run…

A Python toolbox to create adversarial examples that fool neural networks in PyTorch, TensorFlow, and JAX

A security scanner for your LLM agentic workflows

Curated list of backdoor learning papers, surveys, and toolboxes, organizing poisoning-based attacks and defenses in deep learning for researchers…

Standardized adversarial robustness benchmark with a public leaderboard and downloadable model zoo for evaluating ML models against Lp attacks and…