
heretic
Fully automatic censorship removal for language models

Fully automatic censorship removal for language models

The community's most comprehensive, continuously-updated index of research on Large Language Models for software vulnerability detection — papers…

An AI-powered Personal Identifiable Information (PII) scanner.

AI / LLM Red Team Field Manual & Consultant’s Handbook

Syntactic Ghost: An Imperceptible General-purpose Backdoor Attacks on Pre-trained Language Models

🥂 Gracefully face hCaptcha challenge with multimodal large language model.

An alignment auditing agent capable of quickly exploring alignment hypothesis

Automated security analysis pipeline that runs CodeQL queries on GitHub repositories and uses LLMs to classify and filter true vulnerabilities from…

Detects unknown jailbreak attacks in large vision-language models using hidden state analysis and autoencoders, with training and evaluation…

Fully automatic censorship removal for language models

LLM powered fuzzing via OSS-Fuzz.

Reverse Engineering: Decompiling Binary Code with Large Language Models

OGhidra bridges Large Language Models (LLMs) via Ollama with the Ghidra reverse engineering platform, enabling AI-driven binary analysis through…

IDA plugin which queries language models to speed up reverse-engineering

EmailXpose is an open source AI-powered email security system that detects phishing, spam, scams, malware, and social engineering attacks. It goes…

LLM-agent-powered concolic execution engine that instruments source code, summarizes path constraints in natural language, and generates test cases…

Research code and experiments for defending tool-integrated LLM agents against adversarial attacks, extending Agent Security Bench with new defense…