
AMBER-ICI
AMBER ICI v5: local-first Ollama investigative command center with case-scoped evidence, agent chains, hybrid retrieval, streaming analysis, graph…

AMBER ICI v5: local-first Ollama investigative command center with case-scoped evidence, agent chains, hybrid retrieval, streaming analysis, graph…

Open-source AI security benchmarking CLI. Measure how AI models perform offensive security tasks with MITRE ATT&CK analysis and KSM scoring.

Benchmark for evaluating AI agent safety against attacks embedded in skill-facing context, with 155 cases across 6 risk domains, measuring task…

This repository contains the complete record of my three-year research journey, covering the project from foundational concepts to advanced-level…

An OWASP-aligned intentionally vulnerable platform for learning and testing AI, LLM, RAG, MCP, and Agentic AI security.

An empirical security testbed evaluating prompt injection, confused-deputy vulnerabilities, and tool-calling defenses in LLM agents.

DonkAI is a hands-on lab for the OWASP Top 10 for LLM Applications (2025) - no real LLM required.

Official repository for CTFTiny

Hands-on AI security lab platform with 50+ scenarios across prompt injection, agentic system exploitation, model manipulation, and MCP trust boundary…

Security benchmark for evaluating OpenClaw agents against adversarial execution contexts including poisoned files, injected skills, misleading tool…

Open-source prompt injection attack console. Test AI security by firing categorized attacks at any endpoint.

A benchmark for LLM-driven bug discovery: 77 challenges across 43 open-source projects (C/C++/Java).

Explainable security gate for LLM apps — blocks prompt injection with an auditable reason for every decision.

Benchmark measuring AI models' ability to detect vulnerabilities in source code via real bug bounty cases with balanced recall and false-positive…

Code for paper "ActBench: Self-Evolving Benchmark of Behavioral Safety in Cowork Agents"

A small go harness that uses Ollama to orchestrate LLMs in a restricted process flow

A cybersecurity research game measuring how humans detect AI-generated phishing emails. Built as a retro terminal experience.

Voice-based detective interrogation game. Mistral Large 3 + Voxtral STT + ElevenLabs TTS. Built for the Mistral Worldwide Hackathon 2026.