
RLCDAlignBench
Benchmark suite and code for detecting AI alignment failures, with 44 benchmarks across ten failure types and a zero-shot RLCD detector evaluated on…

Benchmark suite and code for detecting AI alignment failures, with 44 benchmarks across ten failure types and a zero-shot RLCD detector evaluated on…

Next-gen logical WAF engine built in SWI-Prolog. Features an inductive learning brain running at 2M+ LIPS with an integrated recursive decoder to…

Enterprise AI agent security toolkit providing pre-flight auditing, configuration hardening, runtime threat detection, and active defense against…

Hybrid machine-learning pipelines for detecting SQL injection in web traffic, combining DistilBERT and BERT-GNN models with adversarial training and…

Detects unknown jailbreak attacks in large vision-language models using hidden state analysis and autoencoders, with training and evaluation…

An Evaluation Agent for Detecting Misinformation and Knowledge Poisoning in Retrieval-Augmented Generation Systems.

Algorithms for outlier, adversarial and drift detection

Real-time cloud-native runtime security agent for Linux that monitors syscalls and container/Kubernetes metadata to detect anomalous behavior and…

Collection of Google Cloud solution examples and operational utilities for audit log monitoring, DLP de-identification, encryption key management,…

Programmable guardrails for LLM chat apps: enforce input/output rails, block jailbreaks and prompt injections, detect hallucination, and mask…

Detects LLM context-leakage attacks by training lightweight behavior probes on log-probabilities, with vLLM offline/server detection pipelines.

Self-referenced local contrast for knowledge-poison detection in retrieval-augmented generation

Simple Anti-cheat library for applications that use C++ on windows. #PastedProtection

Self-hostable AI SOC that fuses security alerts, auto-triages via agentic AI, runs MITRE ATT&CK investigations, and logs every agent decision in a…

Runtime security gateway for AI agents: cryptographically attests tool calls, enforces policies, sandboxes execution, and logs tamper-evident audit…

Analyzes LLM internal states and 100+ attention/probability features to train classifiers that detect document poisoning attacks in RAG systems.

Production AI defense with 7-layer protection: mathematical constraints, object-capability access, distributed O2 consensus, SVETILO ethics. First…

Enumerate various traits from Windows processes as an aid to threat hunting