← Back to brief
Policy & SafetyOfficialPreprintarXiv Cryptography and Security

TYPO: Visual Jailbreaks Expose Safety Gaps in Commercial Image-Generation Models

A new arXiv preprint introduces TYPO, a black-box attack that exploits a safety vulnerability in commercial image-generation models. While these models often block harmful text prompts, TYPO demonstrates that they can be manipulated to generate images containing detailed, readable, and actionable harmful instructions as embedded text. The method outperforms nine prior jailbreak attacks in attack success rate across four commercial models, highlighting a significant gap in current safety alignment.

Why it matters: This work exposes a critical and previously underreported vulnerability in widely used image-generation systems, showing that safety measures for text do not reliably extend to text rendered within images.

Full story at: arXiv Cryptography and Security