
MMSkillRisk
Benchmark and evaluation harness testing whether LLM agents resist malicious instructions hidden in multimodal skill images, with 108 cases across…

Benchmark and evaluation harness testing whether LLM agents resist malicious instructions hidden in multimodal skill images, with 108 cases across…

Browser-local security monorepo with six modules for mobile APK/IPA triage, client-side DAST fuzzing, OSINT directories, offline AI threat scoring,…

Pre-execution action-auditing defense that detects and masks indirect prompt injection in tool-using LLM agents using embedding retrieval and…

EmailXpose is an open source AI-powered email security system that detects phishing, spam, scams, malware, and social engineering attacks. It goes…

Research code for training adapters that make LLM agent backdoors survive benign fine-tuning, with trigger-rate scoring, SWE-bench evaluation, and…

Defense framework that mitigates malicious fine-tuning of LLMs using selective layer recovery and dynamic routing, with training, harmfulness…

CAPTCHA proves you're human. HATCHA proves you're not.

Benchmark measuring AI models' ability to detect vulnerabilities in source code via real bug bounty cases with balanced recall and false-positive…

Project Mantis: Hacking Back the AI-Hacker; Prompt Injection as a Defense Against LLM-driven Cyberattacks

Security benchmark for evaluating OpenClaw agents against adversarial execution contexts including poisoned files, injected skills, misleading tool…

Documents CVE-2026-52618 with a PoC for OS command injection in @webfer/mcp-ansible-drupal via executeDeployment extraVars, plus detection guidance…

B2B software vendor evaluation skill for Claude Code — domain-expert questions, vendor AI agent conversations, evidence-based scoring