Cross-Modal Unlearning Transfer in Vision-Language Models Is Asymmetric and Vulnerable to Typographic Attacks
A systematic study of cross-modal unlearning in vision-language models (LLaVA-1.5, InstructBLIP, IDEFICS) finds that unlearning knowledge in one modality (text or vision) transfers asymmetrically and incompletely to the other. The research shows that typographic attacks—manipulating the visual presentation of text—can recover previously unlearned knowledge, revealing that current unlearning methods are shallow. The proposed CrossInf mitigation strategy reduces the transfer gap by more than half and lowers attack success rates to near zero, while preserving model utility.
Why it matters: This work reveals a critical vulnerability in current unlearning methods for multimodal AI, showing that knowledge can be recovered via cross-modal attacks, and introduces a practical mitigation to improve the safety and reliability of vision-language models.
Full story at: arXiv Cryptography and Security ↗