
MalEval
Official code for the ISSTA 2026 paper: Is "Knowing It’s Malicious" Enough? Evaluating LLMs for Fine-Grained Malware Behavior Auditing

Official code for the ISSTA 2026 paper: Is "Knowing It’s Malicious" Enough? Evaluating LLMs for Fine-Grained Malware Behavior Auditing

Practical black-box adversarial packet generation against encrypted traffic classification with minimal overhead and full packet recoverability.

Causal context attribution and rule-based monitor LLM defense against indirect prompt injection in LLM agents, achieving state-of-art performance on…

Research implementation for mitigating adaptive prompt injections via on-policy distillation, with training recipes and evaluators for SEP, PISmith,…

Code for the paper Certified Unlearning for Neural Networks, ICML 2025

Code for 'Steering the CensorShip: Uncovering Representation Vectors for LLM "Thought" Control'

Implementation of paper "DeeCLIP: A Robust and Generalizable Transformer-Based Framework for Detecting AI-Generated Images"

Automated framework for hijacking safety reasoning in large reasoning models via simulated reasoning traces and iterative prompt refinement to…

This repository provides the official implementation of POISONCRAFT: Practical Poisoning of Retrieval-Augmented Generation for Large Language Models.

The code of VulTriage: Triple-Path Context Augmentation for LLM-Based Vulnerability Detection

Curated reading list and taxonomy of attack and defense research for mobile on-device AI systems, covering adversarial, backdoor, model stealing, and…

Human-evaluated benchmark for assessing LLM performance on real-world vulnerability identification, explanation, and remediation across 15+ languages…

The code for ACM MM2024 (Multimodal Unlearnable Examples: Protecting Data against Multimodal Contrastive Learning)

Inference scaling for LLM safety assurance

Provably secure linguistic steganography embedding secret messages into LLM-generated text via rotation range-coding. Includes embed, extract, and…

Detailed disclosure of CVE-2025-67511, a command injection vulnerability in the CAI framework's SSH tool that allows AI agents to be tricked into…

Research demonstration of indirect prompt injection attacks to control autonomous LLM-based web agents, with tools for trigger optimization and…

Evaluation framework that tests whether large language models follow invisible Unicode-encoded instructions embedded in normal-looking text, with…