RECON: Benchmarking Agent Memory for Compositional Reasoning over Long Contexts
A new benchmark, RECON, evaluates large language model (LLM)-based agents on their ability to perform compositional reasoning over extended contexts of 50,000 to 100,000 tokens in criminal, medical, and financial domains. RECON tests six memory-intensive tasks, including reconstructing multi-hop evidence chains, propagating cascading invalidations, resolving source conflicts, counterfactual reasoning, satisfying temporal constraints, and temporal fact retrieval. The evaluation shows that even the strongest non-Oracle system achieves only 22.4% accuracy, revealing significant limitations in current agent memory and reasoning capabilities.
Why it matters: RECON exposes critical gaps in the ability of state-of-the-art LLM agents to reason over long contexts, which is essential for reliable deployment in complex real-world applications.
Full story at: arXiv AI/ML ↗