
AgentWatcher
Causal context attribution and rule-based monitor LLM defense against indirect prompt injection in LLM agents, achieving state-of-art performance on…

Causal context attribution and rule-based monitor LLM defense against indirect prompt injection in LLM agents, achieving state-of-art performance on…

Human-evaluated benchmark for assessing LLM performance on real-world vulnerability identification, explanation, and remediation across 15+ languages…

Code for 'Steering the CensorShip: Uncovering Representation Vectors for LLM "Thought" Control'

Implementation of paper "DeeCLIP: A Robust and Generalizable Transformer-Based Framework for Detecting AI-Generated Images"

This repository provides the official implementation of POISONCRAFT: Practical Poisoning of Retrieval-Augmented Generation for Large Language Models.

Research implementation for mitigating adaptive prompt injections via on-policy distillation, with training recipes and evaluators for SEP, PISmith,…

Research pipeline for detecting latent indirect prompt-injection exposure signals in agentic LLMs via hidden-state probing, including trace…

Production AI defense with 7-layer protection: mathematical constraints, object-capability access, distributed O2 consensus, SVETILO ethics. First…

Inference scaling for LLM safety assurance

Provably secure linguistic steganography embedding secret messages into LLM-generated text via rotation range-coding. Includes embed, extract, and…

Curated reading list and taxonomy of attack and defense research for mobile on-device AI systems, covering adversarial, backdoor, model stealing, and…

This repository demonstrates a machine learning pipeline for detecting MITRE ATT&CK techniques from logs and enriching the output using a local LLM.

FLNET2023 is a dataset designed for intrusion detection in Federated Learning scenarios. It's built using the CORE emulator, simulating a realistic…

ML-Based behavioral endpoint detection system for Linux machines

Detects LLM context-leakage attacks by training lightweight behavior probes on log-probabilities, with vLLM offline/server detection pipelines.

Stop prompt injection attacks before they reach your LLM — zero API costs, runs entirely locally, integrates in 2 minutes. Prompt injection is the…

ML-based detection of Zombie ZIP archive header evasion attacks (CVE-2026-0866)

Adversary-resilient deep learning architecture for secure 5G indoor localization, combining CNN and multi-head attention to defend against signal…