← Back to brief
ResearchOfficialPreprintarXiv Machine Learning

The Information Shadow: Structural Limits on What Language Models Can Learn

A new preprint introduces the concept of the 'information shadow'—a set of phenomena that text-trained language models cannot learn, regardless of model scale. The authors formally identify three types of structural limits: (1) information that language cannot express, (2) functions that are statistically non-identifiable from training data, and (3) functions that are representable but unreachable by gradient-based training. Each limit is demonstrated with provable probes and controls that rule out capacity or modality artifacts.

Why it matters: This work provides a formal framework for understanding fundamental, scale-independent limits of language models, informing benchmark design and capability auditing.

Full story at: arXiv Machine Learning