
promptfoo
Test your prompts, agents, and RAGs. Red teaming/pentesting/vulnerability scanning for AI. Compare performance of GPT, Claude, Gemini, DeepSeek, and…

Test your prompts, agents, and RAGs. Red teaming/pentesting/vulnerability scanning for AI. Compare performance of GPT, Claude, Gemini, DeepSeek, and…

Self-hosted multi-agent environment for Go with LLM-powered pentesting agents (exploiter, reverser, threathunter, webscanner) that automate…

Modular LLM vulnerability scanner that probes for hallucination, data leakage, prompt injection, jailbreaks, and toxicity using static, dynamic, and…

Suppress vulnerabilities applying Kubernetes context to scans

OGhidra bridges Large Language Models (LLMs) via Ollama with the Ghidra reverse engineering platform, enabling AI-driven binary analysis through…

C-based Android static analysis framework for decompilation, secret detection, endpoint discovery, permission analysis, and native library scanning…

🐢 Open-Source Evaluation & Testing library for LLM Agents

Curated reading list and taxonomy of attack and defense research for mobile on-device AI systems, covering adversarial, backdoor, model stealing, and…

A fast universal code security scanner, written in Rust. Batteries included: supports 14 languages, TUI for triage, secrets, post-quantum audits,…

Fully automatic censorship removal for language models

ML-based detection of Zombie ZIP archive header evasion attacks (CVE-2026-0866)

Reconstructs legacy Windows binaries into C source by pairing Ghidra decompiler exports with local LLMs, producing compile-checked candidates and…

An autonomous red-teaming engine for LLMs. RedThread manages the full security lifecycle: generating adversarial attacks, executing precision…

Detects process injection and memory manipulation used by malware. Finds RWX regions, shellcode patterns, API hooks, thread hijacking, and process…

Security Scanner for Agent Skills

A diagnostic framework for measuring LLM vulnerability to Affective Contextual Erosion (ACE) and related liminal attack vectors. **Delirium** is not…

A machine learning tool that ranks strings based on their relevance for malware analysis.

reverse engineering Gemini's SynthID detection