
autoguardrails
Alignment-research scaffold (autoresearch-style) for LLM guardrails: search over a single policy.md surface

Alignment-research scaffold (autoresearch-style) for LLM guardrails: search over a single policy.md surface

Attempt at Obfuscated version of SharpCollection

Delving into the Realm of LLM Security: An Exploration of Offensive and Defensive Tools, Unveiling Their Present Capabilities.

See adversary, do adversary: Simple execution of commands for defensive tuning/research (now with more ELF on the shelf)

An autonomous red-teaming engine for LLMs. RedThread manages the full security lifecycle: generating adversarial attacks, executing precision…

Demonstrates using machine learning to predict random number generator sequences, highlighting cryptographic weaknesses through adversarial analysis.

Open-source cross-modal and multimodal prompt injection test suite. 250,000+ attack payloads across text, image, document, and audio modalities.…

Proof-of-concept exploit that bypasses ASLR on Intel CPUs by abusing branch target buffer and speculative execution to leak randomized addresses via…

Open-source AI security platform providing perimeter defense for LLMs and AI agents through swarm analysis, policy enforcement, adversarial testing,…

[ICLR 2026] - Official repo for the paper: "RedBench: A Universal Dataset for Comprehensive Red Teaming of Large Language Models"

A high-severity prompt injection flaw in Claude AI proves that even the smartest language models can be turned into weapons — all with a few lines of…

A Python library for Secure and Explainable Machine Learning Documentation available @ https://secml.gitlab.io Follow us on Twitter @…

Framework for auditing machine learning algorithms against adversarial attacks and biases, providing educational tools and an academic paper to…

Reverse Shell Detection with Machine Learning

Practical black-box adversarial packet generation against encrypted traffic classification with minimal overhead and full packet recoverability.

Causal context attribution and rule-based monitor LLM defense against indirect prompt injection in LLM agents, achieving state-of-art performance on…

A prompt injection in a code‑review bot that executes AI‑generated fixes in a sandbox. The sandbox uses a blacklist to prevent dangerous commands,…