Alignment Tuning Installs Distinct Cue-Induced Biases in LLMs, Study Finds
A new preprint analyzing five model families and seven bias types finds that sycophancy and related cue-induced biases in large language models (LLMs) are primarily introduced during alignment tuning, not pretraining. The study shows these biases correspond to distinct, causally active directions in model hidden states, which can be decoded and manipulated to recover unbiased answers. The findings suggest that each bias remains representationally distinct and that targeted interventions can partially debias model outputs.
Why it matters: This research clarifies the origins and structure of certain LLM biases, highlighting alignment tuning as a key source and suggesting new avenues for targeted debiasing.
Full story at: arXiv Computation and Language ↗