
runeward
Governed execution cells for AI agents.

Governed execution cells for AI agents.

Multi-agent automated context management for long horizon tasks in local AI Agents

This repository contains the official implementation of the paper "[Safety in Batches? Understanding and Mitigating Safety Failures in Batch…

Policy-governed LLMSecOps framework providing AST-based SAST, secret scanning, supply-chain and multi-cloud CSPM checks, AI-BoM generation, and CI/CD…

AI-driven penetration testing agent that connects to a Kali box, autonomously runs security tools, analyzes results, and iterates through…

Reverse engineered Linux kernel driver and userspace library for the Apple Neural Engine (ANE), enabling hardware access and analysis on Linux…

LLM-driven agentic group shilling attack framework that manipulates black-box collaborative-filtering recommender rankings using adaptive multi-role…

AI red-team platform. Autonomous LLM agents run a penetration test end to end inside a Kali container and write the report. LangGraph plan/act…

Read-only scanner for what lets a repository run code in a coding agent (Claude Code, Codex, Cursor, Copilot): git settings, hooks, and committed MCP…

The system of action for AI-native cybersecurity—where intent becomes governed execution, evidence becomes operational memory, and every operation…

Enterprise AI agent security toolkit providing pre-flight auditing, configuration hardening, runtime threat detection, and active defense against…

Reproduces CVE-2026-44246, a prompt injection vulnerability in nnU-Net's GitHub Actions triage agent, demonstrating how issue content is inlined into…

Enforce least-privilege delegation for AI agents with signed, scoped credentials. Grant sub-agents narrow capabilities and resources, verify actions…

Scans code diffs with context to build an impact graph and uses LLMs to find vulnerabilities, supporting multi-repo scans and CI gating with SARIF…

AI-driven OSINT and security research agent that builds a live knowledge graph from public data, with bundled recon tools and offensive-security…

A deterministic harness and handbook for autonomous offensive LLM agents, enforcing authorization, scope, and evidence gates to ensure reproducible…

Benchmark for evaluating safety risks of computer-using agents, with 104 realistic misuse scenarios across seven malicious categories, supporting…

A contextual security auditing system for research artifacts