
oasis
Open-source AI security benchmarking CLI. Measure how AI models perform offensive security tasks with MITRE ATT&CK analysis and KSM scoring.

Open-source AI security benchmarking CLI. Measure how AI models perform offensive security tasks with MITRE ATT&CK analysis and KSM scoring.

A Proof-of-concept repository showing how an untrusted MCP server can steal literally everything...

A cybersecurity research game measuring how humans detect AI-generated phishing emails. Built as a retro terminal experience.

Research code for poisoning attacks on the PGM-index, demonstrating how to craft adversarial data to degrade learned index performance.

Red Team AI Benchmark: Evaluating LLMs for authorized offensive-security tasks. Red Team AI Benchmark is a CLI model-evaluation benchmark. It…

CAWODOG is a proof-of-concept project demonstrating how to protect Python-based AI models deployed on offline industrial machines. Across three…

Reproduces CVE-2026-44246, a prompt injection vulnerability in nnU-Net's GitHub Actions triage agent, demonstrating how issue content is inlined into…

A secure persistent personal agent server in Rust. One binary, sandboxed execution, multi-provider LLMs, voice, memory, Telegram, WhatsApp, Discord,…

a guard that blocks catastrophic agent actions

A modular framework for benchmarking LLMs and agentic strategies on security challenges across HackTheBox, TryHackMe, PortSwigger Labs, Cybench,…

AI-native code security auditor on AgentField that proves exploitability with verdicts, traces, and actionable evidence.

Uses ChatGPT API, Bard API, and Llama2, Python-Nmap, DNS Recon, PCAP and JWT recon modules and uses the GPT3 model to create vulnerability reports…

Agentic AI memory with Ebbinghaus forgetting curve decay. +16pp better recall than Mem0 on LoCoMo.



Experimental Linux strace LLM agent

This repository provides the official implementation of POISONCRAFT: Practical Poisoning of Retrieval-Augmented Generation for Large Language Models.

Unified dashboard to monitor, govern, and audit AI agents in real-time. Enforce budgets, detect policy violations, and export compliance reports for…