
IF-Guide
[NeurIPS '25] Code for Paper "IF-Guide: Influence Function-Guided Suppression of Harmful Training Data for Reducing LLM Toxicity"

[NeurIPS '25] Code for Paper "IF-Guide: Influence Function-Guided Suppression of Harmful Training Data for Reducing LLM Toxicity"

Code for paper "ActBench: Self-Evolving Benchmark of Behavioral Safety in Cowork Agents"

Code for the paper Certified Unlearning for Neural Networks, ICML 2025

This repository provides the official implementation of POISONCRAFT: Practical Poisoning of Retrieval-Augmented Generation for Large Language Models.

Inference scaling for LLM safety assurance

Proof-of-concept for CVE-2026-23001: demonstrates RCE through unsafe pickle deserialization in Hugging Face Transformers by crafting a malicious…

Detailed disclosure of CVE-2025-67511, a command injection vulnerability in the CAI framework's SSH tool that allows AI agents to be tricked into…

Security & Compliance bodyguard for OpenClaw agents

The code for ACM MM2024 (Multimodal Unlearnable Examples: Protecting Data against Multimodal Contrastive Learning)

A Proof-of-concept repository showing how an untrusted MCP server can steal literally everything...

Production AI defense with 7-layer protection: mathematical constraints, object-capability access, distributed O2 consensus, SVETILO ethics. First…

Voice-based detective interrogation game. Mistral Large 3 + Voxtral STT + ElevenLabs TTS. Built for the Mistral Worldwide Hackathon 2026.

Silent dependency injection through AI documentation pipelines. 240 isolated Docker runs proving Context Hub's zero-sanitization MCP server lets…

FastGPT Python sandbox escape chain audit tool (CVE-2026-32128 related, v4.14.8 inspect chain)

Security Advisory: Stored Cross-Site Scripting Via Agent Messages Leading To Session Token Theft (openclaw-dashboard)

Detects LLM context-leakage attacks by training lightweight behavior probes on log-probabilities, with vLLM offline/server detection pipelines.

CVE-2026-67598 — Emlog Pro: disabled TLS certificate validation in AI assistant (MITM → API-key theft). CWE-295, CVSS 9.1. Reported by @IlhomjonR.

CVE-2026-45033 PoC for Claude Code, not Github Copilot. Worked for Haiku 4.5.