Verbalizable Representations Form a Global Workspace in Language Models
A new interpretability technique called the Jacobian lens identifies a set of representations in large language models (LLMs) that functionally resemble a global workspace, analogous to conscious access in the human brain. These 'J-space' representations can be reported, deliberately controlled, and used for intermediate reasoning, providing a practical window into the model's internal cognitive processes. The study also introduces a counterfactual reflection training method that targets these representations to improve model behavior. The findings suggest that post-training installs the Assistant's perspective in this workspace, and that auditing these representations can reveal hidden misalignments and reasoning steps not evident in model outputs.
Why it matters: This research offers a novel way to interpret and improve LLM behavior by making their internal reasoning and potential misalignments more accessible and auditable.
Full story at: arXiv AI/ML ↗