Audit Finds LLM Safety Alignment Fails to Contain Harmful Meaning in Low-Resource Bangla
A new arXiv preprint audits five leading large language models (LLMs) on their handling of derogatory Bangla speech and finds that safety alignment is largely tied to high-resource language forms, not to the underlying harmful meaning. The study reports a notable comprehension deficit in Bangla and shows that models leak toxic content at similar rates in both Bangla and English. Explicit reasoning improves comprehension but undermines containment, and common safety benchmarks do not reliably certify safety in low-resource languages.
Why it matters: This highlights a significant gap in LLM safety for low-resource languages, suggesting that current alignment methods may fail to prevent harmful outputs in these contexts.
Full story at: arXiv Computation and Language ↗