Semantic Primes as Explanatory Primitives for Emotion in Large Language Models
A new preprint proposes using semantic primes from the Natural Semantic Metalanguage (NSM) as foundational elements to explain and control emotions in large language models (LLMs). Experiments on four instruction-tuned LLMs demonstrate that NSM primes are more recoverable, controllable, and selective than appraisal-based directions, with interventions on NSM primes controlling emotion about three times as strongly and twice as selectively. The study suggests that NSM primes may serve as more effective explanatory primitives for emotion in LLMs than existing approaches.
Why it matters: This work introduces a potentially more interpretable and effective framework for understanding and manipulating emotions in LLMs, which could impact model alignment and safety.
Full story at: arXiv AI/ML ↗