
D-Scan
Analyzes LLM internal states and 100+ attention/probability features to train classifiers that detect document poisoning attacks in RAG systems.
ai-securityanomaly-detectiondefensive-tools+1
3

Analyzes LLM internal states and 100+ attention/probability features to train classifiers that detect document poisoning attacks in RAG systems.