
Defenses-for-Tool-Integrated-LLM
Research code and experiments for defending tool-integrated LLM agents against adversarial attacks, extending Agent Security Bench with new defense…

Research code and experiments for defending tool-integrated LLM agents against adversarial attacks, extending Agent Security Bench with new defense…

Syntactic Ghost: An Imperceptible General-purpose Backdoor Attacks on Pre-trained Language Models

Research code for extracting and training safety-awareness directions in multimodal LLMs to improve refusal behavior while limiting benign-task drift.

Detects unknown jailbreak attacks in large vision-language models using hidden state analysis and autoencoders, with training and evaluation…

🥂 Gracefully face hCaptcha challenge with multimodal large language model.

LLM powered fuzzing via OSS-Fuzz.

The community's most comprehensive, continuously-updated index of research on Large Language Models for software vulnerability detection — papers…

Reverse Engineering: Decompiling Binary Code with Large Language Models

IDA plugin which queries language models to speed up reverse-engineering

A productionized greedy coordinate gradient (GCG) attack tool for large language models (LLMs)

This repository provides the official implementation of POISONCRAFT: Practical Poisoning of Retrieval-Augmented Generation for Large Language Models.

Ensemble framework for software vulnerability detection and repair using multiple large language models, with consensus analysis and evaluation tools…

A novel adversarial attack on LLM based on the Exponentiated Gradient Descent technique.

A diagnostic framework for measuring LLM vulnerability to Affective Contextual Erosion (ACE) and related liminal attack vectors. **Delirium** is not…

Code for ACL 2026 (main) paper "DeepGuard: Secure Code Generation via Multi-Layer Semantic Aggregation"

The code of VulTriage: Triple-Path Context Augmentation for LLM-Based Vulnerability Detection

Fully automatic censorship removal for language models

EmailXpose is an open source AI-powered email security system that detects phishing, spam, scams, malware, and social engineering attacks. It goes…