
ClearML-CVE-2024-24590
Proof of concept for CVE-2024-24590

Proof of concept for CVE-2024-24590

Collection of CVE(work) on tenserflow binary pwning it

A cybersecurity research game measuring how humans detect AI-generated phishing emails. Built as a retro terminal experience.

Research code for poisoning attacks on the PGM-index, demonstrating how to craft adversarial data to degrade learned index performance.

a guard that blocks catastrophic agent actions

A secure persistent personal agent server in Rust. One binary, sandboxed execution, multi-provider LLMs, voice, memory, Telegram, WhatsApp, Discord,…

Uses ChatGPT API, Bard API, and Llama2, Python-Nmap, DNS Recon, PCAP and JWT recon modules and uses the GPT3 model to create vulnerability reports…

A Proof-of-concept repository showing how an untrusted MCP server can steal literally everything...

CAWODOG is a proof-of-concept project demonstrating how to protect Python-based AI models deployed on offline industrial machines. Across three…

A modular framework for benchmarking LLMs and agentic strategies on security challenges across HackTheBox, TryHackMe, PortSwigger Labs, Cybench,…

Agentic AI memory with Ebbinghaus forgetting curve decay. +16pp better recall than Mem0 on LoCoMo.

AI-native code security auditor on AgentField that proves exploitability with verdicts, traces, and actionable evidence.


Red Team AI Benchmark: Evaluating LLMs for authorized offensive-security tasks. Red Team AI Benchmark is a CLI model-evaluation benchmark. It…

Artefacts for blog post on finding CVE-2025-37899 with o3

Experimental Linux strace LLM agent

Open-source AI security benchmarking CLI. Measure how AI models perform offensive security tasks with MITRE ATT&CK analysis and KSM scoring.