← Back to brief
Policy & SafetyOfficialPreprintarXiv Cryptography and Security

HumorSafe: Benchmarking Latent Safety Risks in LLM-Driven Content Humorization

A new preprint reveals that using humor as an indirect refusal mechanism in large language models (LLMs) can introduce latent safety risks, such as stereotypes and toxicity. The authors introduce HumorSafe, a framework for evaluating these risks, and HumorPIA, a prompt injection attack that exploits humor-based defenses to covertly increase toxicity while maintaining a high apparent safety rate.

Why it matters: This research exposes a previously overlooked vulnerability in LLM safety mechanisms, showing that humor-based defenses can covertly propagate harmful content and evade current detection methods.

Full story at: arXiv Cryptography and Security