← Back to brief
ResearchOfficialPreprintarXiv Computation and Language

KyrgyzLLM-Bench: First Systematic Benchmark for Kyrgyz Language Understanding

Researchers introduce KyrgyzLLM-Bench, a benchmark suite for evaluating large language models (LLMs) in Kyrgyz. The suite includes two natively authored datasets (KyrgyzMMLU and KyrgyzRC) and translated, manually post-edited versions of WinoGrande, HellaSwag, BoolQ, and TruthfulQA. Evaluation of 26 models shows that cross-lingual transfer varies by task, with translation artifacts notably affecting HellaSwag performance. Few-shot prompting improves some open-source models on reading comprehension, but results are inconsistent for proprietary models on translated tasks.

Why it matters: This work provides the first large-scale, natively authored evaluation benchmark for Kyrgyz, enabling more accurate assessment of LLM capabilities and cross-lingual transfer in a low-resource language.

Full story at: arXiv Computation and Language