ADAGE: A Language-Agnostic Pipeline for Analogical Reasoning Evaluation
A new arXiv preprint introduces ADAGE, a pipeline designed to create translation-free benchmarks for evaluating analogical reasoning in multiple languages. By combining native-speaker curation with LLM-assisted generation, the authors developed benchmarks for Arabic, Amharic, and Japanese, revealing that large language models show a significant drop in accuracy—12 to 52 percentage points—on these culturally-grounded tasks compared to English. The pipeline and benchmarks are publicly released.
Why it matters: This work exposes a substantial gap in multilingual AI reasoning capabilities and provides new tools to measure and address culturally-specific reasoning challenges.
Full story at: arXiv Computation and Language ↗