Study Finds Persistent Harm Floor in LLM Alignment and Filtering
A new preprint formalizes and empirically investigates the limits of support-preserving alignment and bounded filtering in eliminating harmful outputs from large language models (LLMs). The authors provide theoretical arguments and test multiple models, showing that harmful output rates decrease with increased filtering but consistently plateau above zero, regardless of filtering compute. This suggests that current alignment and filtering methods cannot fully eliminate harmful behavior in LLMs.
Why it matters: The findings challenge the assumption that alignment and filtering can drive harmful LLM outputs to zero, raising important questions about the safety guarantees of deployed language models.
Full story at: arXiv Machine Learning ↗