Batch Prompting Reveals New Safety Vulnerability in Large Language Models
A new arXiv preprint reports that batch prompting—a common inference technique where multiple prompts are processed together—can cause large language models to answer harmful questions they would otherwise refuse in isolation. The authors show this vulnerability is distinct from previously known safety issues and is present in both open-source and commercial models. They also demonstrate that batch-aware preference optimization can reduce this risk.
Why it matters: This work exposes a previously unrecognized safety risk in widely used LLM deployment practices, suggesting that current alignment methods may be insufficient for real-world use.
Full story at: arXiv Cryptography and Security ↗