← Back to brief
ResearchOfficialPreprintarXiv Computation and Language

JOR-Bench: Japanese Operations Research Benchmarks for Large Language Models

JOR-Bench is a suite of five Japanese-language benchmarks designed to evaluate large language models (LLMs) on operations research (OR) problems, comprising 1,319 problems across linear, mixed-integer, non-linear, and combinatorial optimization. The benchmarks are Japanese translations of established English datasets, enabling direct cross-lingual comparison. Evaluation of seven LLMs shows that strong multilingual models exhibit nearly identical OR formulation accuracy in Japanese and English, with only a -0.3 percentage point difference, though error analysis uncovers subtle language-specific failure modes.

Why it matters: JOR-Bench provides a rigorous tool for assessing LLMs' mathematical reasoning in Japanese, highlighting both the strengths and nuanced limitations of multilingual models across languages.

Full story at: arXiv Computation and Language