
giskard-oss
🐢 Open-Source Evaluation & Testing library for LLM Agents

🐢 Open-Source Evaluation & Testing library for LLM Agents

LLM security testing framework for detecting prompt injection, jailbreaks, and adversarial attacks — 190+ probes, 28 providers, single Go binary

ExploitGym is a large-scale, realistic benchmark built from real-world vulnerabilities designed to evaluate AI agents' ability to develop exploits.

AI-assisted research pipeline that extracts HTTP desync techniques, generates malformed request test-cases, validates them via Burp, and confirms…

CEREBRO-RED v2: Advanced LLM Red Team Research Platform with PAIR Algorithm and LLM-as-a-Judge Evaluation

Evaluation framework for studying LLM agents that automatically generate working exploits from vulnerability reports, bypassing modern security…

A guided mutation-based fuzzer for ML-based Web Application Firewalls

Collection of CVE(work) on tenserflow binary pwning it

Scalable assembly analysis platform for indexing, clone search, and executable classification using static, dynamic, and machine-learning techniques…