MeetingToM: Benchmarking Multimodal LLMs on Theory-of-Mind in Multi-Party Meetings
Researchers have introduced MeetingToM, a new benchmark designed to evaluate multimodal large language models (MLLMs) on theory-of-mind reasoning within multi-party meetings. The benchmark addresses complex social phenomena such as pseudo-consensus—where apparent agreement conceals private dissent—and assesses models on mental state prediction, addressee understanding, and group consensus reasoning. Initial analyses show that current MLLMs face significant challenges in integrating non-verbal cues and inferring hidden attitudes.
Why it matters: MeetingToM exposes key limitations in current multimodal LLMs' ability to understand nuanced social dynamics, which is crucial for developing more human-like AI systems for real-world group interactions.
Full story at: arXiv Computation and Language ↗