
safety-awareness
Research code for extracting and training safety-awareness directions in multimodal LLMs to improve refusal behavior while limiting benign-task drift.
ai-securityeducationmachine-learning+1

Research code for extracting and training safety-awareness directions in multimodal LLMs to improve refusal behavior while limiting benign-task drift.

This repository includes the source code used in the "Characterization and Detection of Cross-Router Covert Channels" paper.