
ActGuard
Pre-execution action-auditing defense that detects and masks indirect prompt injection in tool-using LLM agents using embedding retrieval and…

Pre-execution action-auditing defense that detects and masks indirect prompt injection in tool-using LLM agents using embedding retrieval and…

Fix-Like Artifacts With Embedded Defects

GNU Radio module for physical layer security using deep learning channel fingerprinting, feature quantization, and SHA3-512 key generation from SDR…

Research code and experiments for defending tool-integrated LLM agents against adversarial attacks, extending Agent Security Bench with new defense…

Research code reproducing multi-turn LLM jailbreak experiments (FITD, MRCJ, ActorAttack, X-Teaming) from the SoK intent-oriented systematization…

Research code for Rubric-Induced Preference Drift (RIPD): evolutionary rubric search, benchmark-preserving selection, and DPO policy misalignment…

Multi-agent automated context management for long horizon tasks in local AI Agents

Open-source threat intelligence platform for malware and observable analysis. Enriches IPs, domains, URLs, and hashes with external sources, performs…

Robust audio watermarking framework embedding binary messages into magnitude spectrograms, with differentiable attack simulation, Q-Former pooling,…

Research code for poisoning attacks on the PGM-index, demonstrating how to craft adversarial data to degrade learned index performance.

Hybrid machine-learning pipelines for detecting SQL injection in web traffic, combining DistilBERT and BERT-GNN models with adversarial training and…

Trajectory-aware evolutionary search framework for red-teaming LLM agents over MCP servers, generating adversarial prompts to map vulnerability…

Reference implementation of a multi-bit LLM watermarking scheme using coded payload spreading, unbiased reweighting, and soft-decision ECC decoding…

Research code for extracting and training safety-awareness directions in multimodal LLMs to improve refusal behavior while limiting benign-task drift.

Black-box attack framework that hijacks reasoning in agentic retrieval-augmented generation systems by injecting poisoned documents, with support for…

Detects unknown jailbreak attacks in large vision-language models using hidden state analysis and autoencoders, with training and evaluation…

Research code for detecting and detoxifying backdoors in text-to-image diffusion models, with pipelines for Stable Diffusion v1.4, v1.5, XL, and 3…

Research pipeline for detecting latent indirect prompt-injection exposure signals in agentic LLMs via hidden-state probing, including trace…