
heretic
Fully automatic censorship removal for language models

Fully automatic censorship removal for language models

Research pipeline for detecting latent indirect prompt-injection exposure signals in agentic LLMs via hidden-state probing, including trace…

Re-play Security Events

An Open-Source Package for Textual Adversarial Attack.

Framework for auditing machine learning algorithms against adversarial attacks and biases, providing educational tools and an academic paper to…

Research implementation for mitigating adaptive prompt injections via on-policy distillation, with training recipes and evaluators for SEP, PISmith,…

Curated list of backdoor learning papers, surveys, and toolboxes, organizing poisoning-based attacks and defenses in deep learning for researchers…

The agent that grows with you

A curated list of useful resources that cover Offensive AI.

A list of covert channels and steganography/steganalysis resources (books, papers & tools)

Your own personal AI assistant. Any OS. Any Platform. The lobster way. 🦞

List of tools & datasets for anomaly detection on time-series data.

Test your prompts, agents, and RAGs. Red teaming/pentesting/vulnerability scanning for AI. Compare performance of GPT, Claude, Gemini, DeepSeek, and…

Interactive dashboards and libraries for responsible AI model debugging, covering error analysis, fairness, interpretability, counterfactuals, causal…

Interpretability and explainability of data and machine learning models

ESPectre - Motion detection system based on Wi-Fi spectre analysis (CSI), with Home Assistant integration.

A LSTM based framework for handling multiclass imbalance in DGA botnet detection
