#1LLM security, prompt injection, model extraction, adversarial AI, and AI red teaming tools.
Kitploit recommended

Adversary-resilient deep learning architecture for secure 5G indoor localization, combining CNN and multi-head attention to defend against signal…
Research demonstration of indirect prompt injection attacks to control autonomous LLM-based web agents, with tools for trigger optimization and…

Code for the paper Certified Unlearning for Neural Networks, ICML 2025

The code for ACM MM2024 (Multimodal Unlearnable Examples: Protecting Data against Multimodal Contrastive Learning)

StealthRL: RL framework for adversarially paraphrasing AI text to stress-test detector robustness.

[NeurIPS 2025] An official source code for paper "GuardReasoner-VL: Safeguarding VLMs via Reinforced Reasoning".

[NeurIPS '25] Code for Paper "IF-Guide: Influence Function-Guided Suppression of Harmful Training Data for Reducing LLM Toxicity"

Pre-install security for AI agents, npm packages, and MCP servers. Zero-dep local static analysis; normal scans never execute package code.


A diagnostic framework for measuring LLM vulnerability to Affective Contextual Erosion (ACE) and related liminal attack vectors. **Delirium** is not…

Autonomous AI pentesting engine, continuous offensive security across web, cloud, AD & Kubernetes. Agentic reasoning + real exploit execution deliver…

AI Smart Contract Security Analysis and PoC Generation Framework

AIRecon is an autonomous cybersecurity agent that combines a self-hosted Large Language Model (Ollama) with a Kali Linux Docker sandbox and a Textual…

Provably secure linguistic steganography embedding secret messages into LLM-generated text via rotation range-coding. Includes embed, extract, and…

Evaluation framework for AI penetration testing agents that measures validated vulnerability discovery using LLM-based semantic matching, bipartite…

CVE-2026-40487

AI-powered offensive security testing using autonomous agents, directly in your terminal.

ExploitGym is a large-scale, realistic benchmark built from real-world vulnerabilities designed to evaluate AI agents' ability to develop exploits.