← Back to brief
ResearchOfficialPreprintarXiv Computation and Language

Relay-Bench: New Benchmark Tests LLMs on Multi-Domain Reasoning Chains

Researchers introduce Relay-Bench, a text-only benchmark designed to evaluate large language models (LLMs) on their ability to solve multi-domain reasoning chains. The benchmark features composite problems that combine subproblems from distinct domains such as visual reasoning, coding, math, information extraction, and data analysis. The top-performing model, GPT-5.5 (xHigh), achieves a score of 43.3%, indicating substantial room for improvement. Models are allowed to use external tools like code execution and web search during evaluation.

Why it matters: Relay-Bench offers a more comprehensive and challenging assessment of LLMs' capacity to integrate and reason across multiple domains, revealing current limitations in multi-step reasoning.

Full story at: arXiv Computation and Language