Simpler Attack Pipeline Outperforms Complex Methods on Vision-Language Models
Researchers introduce SimVLA, a simplified adversarial attack pipeline for vision-language models that surpasses state-of-the-art baselines in transferability while requiring less computational time and memory. The study identifies issues in existing complex pipelines, such as inappropriate cross-modal interactions and excessive operations, and demonstrates that a streamlined approach can be more effective. Experiments across multiple datasets and tasks show SimVLA's superior performance and efficiency.
Why it matters: This work suggests that simpler, domain-informed adversarial attack methods can outperform more complex approaches, informing both attack and defense strategies for vision-language models.
Full story at: arXiv Cryptography and Security ↗