
promptfoo
Test your prompts, agents, and RAGs. Red teaming/pentesting/vulnerability scanning for AI. Compare performance of GPT, Claude, Gemini, DeepSeek, and…

Test your prompts, agents, and RAGs. Red teaming/pentesting/vulnerability scanning for AI. Compare performance of GPT, Claude, Gemini, DeepSeek, and…

Dynamic and static analysis with Real Time Malware Analysis with Antivirus for Windows, including open-source XDR (3 EDR projects), ClamAV, YARA-X,…

Security scanner for AI/ML model files. Detects malicious code, backdoors, and vulnerabilities before deployment

YC (S26) | Open Computer History | Continuously record your company computer work, map your workflows, help you find work worth automating, and power…

Real-time global intelligence dashboard. AI-powered news aggregation, geopolitical monitoring, and infrastructure tracking in a unified situational…

Self-hostable AI SOC that fuses security alerts, auto-triages via agentic AI, runs MITRE ATT&CK investigations, and logs every agent decision in a…

Security Scanner for Agent Skills

A curated list of useful resources that cover Offensive AI.

Modular LLM vulnerability scanner that probes for hallucination, data leakage, prompt injection, jailbreaks, and toxicity using static, dynamic, and…

A serverless networking protocol designed for resilient state synchronization between autonomous agents in fragmented, low-bandwidth networks

Benchmark harness measuring where prompt injection defenses fire in tool-using LLM agent pipelines, tracking canary tokens across exposed, persisted,…

Benchmark and evaluation harness testing whether LLM agents resist malicious instructions hidden in multimodal skill images, with 108 cases across…

Semantic search over videos using Gemini Embedding 2 or Qwen3-VL.

A fast universal code security scanner, written in Rust. Batteries included: supports 14 languages, TUI for triage, secrets, post-quantum audits,…

Python framework for building LLM workflows as state machines with formal verification via Z3 theorem proving, CTL model checking, and conformal…

The world's fastest agentic crawler. Reclaimed. Reinvented. Ready for war.

Open-source OCR engine with LSTM neural network models for extracting text from images and scanned documents in 100+ languages via CLI, C/C++…

Research code for HARDE, an agent harness that probes and adaptively optimizes components for runtime risk detection and execution control across…