
refusal-relocates
Refusal localizes, the damage relocates, safety layers under few-sample fine-tuning

Refusal localizes, the damage relocates, safety layers under few-sample fine-tuning

Open-weight LLMs and training/evaluation code for defending against prompt injection attacks, with benchmarks for agentic tool-calling and…

0-day malware detection for binaries, source & scripts (that doesn't suck)

DNS Proxy that is simple and fast with not so simple features. Focused on routed DNS forwarding, filtering and parental control.

Benchmarking prompt injection detections for web agents.

Measuring Open Privilege in Agent Defenses

Proof-of-concept that poisons MLflow registered models via the REST API, embedding a malicious pickle to trigger RCE when the model is loaded.

Bidirectional token-classification model for PII detection and masking in text, with CLI for redaction, evaluation, and finetuning on-premises.

Fix-Like Artifacts With Embedded Defects

Research code and experiments for defending tool-integrated LLM agents against adversarial attacks, extending Agent Security Bench with new defense…

Open-source threat intelligence platform for malware and observable analysis. Enriches IPs, domains, URLs, and hashes with external sources, performs…

Hybrid machine-learning pipelines for detecting SQL injection in web traffic, combining DistilBERT and BERT-GNN models with adversarial training and…

Research pipeline for detecting latent indirect prompt-injection exposure signals in agentic LLMs via hidden-state probing, including trace…

A list of covert channels and steganography/steganalysis resources (books, papers & tools)

An adversarial example library for constructing attacks, building defenses, and benchmarking both

Ready-to-use OCR with 80+ supported languages and all popular writing scripts including Latin, Chinese, Arabic, Devanagari, Cyrillic and etc.

Curated list of backdoor learning papers, surveys, and toolboxes, organizing poisoning-based attacks and defenses in deep learning for researchers…

List of tools & datasets for anomaly detection on time-series data.