
heretic
Fully automatic censorship removal for language models

Fully automatic censorship removal for language models

Framework for auditing machine learning algorithms against adversarial attacks and biases, providing educational tools and an academic paper to…

Python library for adversarial machine learning security, enabling red and blue teams to run evasion, poisoning, extraction, and inference attacks…

The agent that grows with you

A Python pickling decompiler and static analyzer

Curated list of backdoor learning papers, surveys, and toolboxes, organizing poisoning-based attacks and defenses in deep learning for researchers…

Research pipeline for detecting latent indirect prompt-injection exposure signals in agentic LLMs via hidden-state probing, including trace…

A privacy-first app that strips AI watermarks from content you own.

An Open-Source Package for Textual Adversarial Attack.

Real-time global intelligence dashboard. AI-powered news aggregation, geopolitical monitoring, and infrastructure tracking in a unified situational…

AI / LLM Red Team Field Manual & Consultant’s Handbook

An autonomous red-teaming engine for LLMs. RedThread manages the full security lifecycle: generating adversarial attacks, executing precision…

Research implementation for mitigating adaptive prompt injections via on-policy distillation, with training recipes and evaluators for SEP, PISmith,…

🐢 Open-Source Evaluation & Testing library for LLM Agents

Automated security analysis pipeline that runs CodeQL queries on GitHub repositories and uses LLMs to classify and filter true vulnerabilities from…

A structured knowledge base covering AI security fundamentals, threat modeling, red team offensive techniques, and blue team defenses, including LLM…

NVR with realtime local object detection for IP cameras

Research code for detecting and detoxifying backdoors in text-to-image diffusion models, with pipelines for Stable Diffusion v1.4, v1.5, XL, and 3…