← Back to brief
ResearchOfficialPreprintarXiv Computation and Language

Pancasila-Dilemmas: New Benchmark Reveals LLMs Struggle with Indonesian Values

Researchers have introduced Pancasila-Dilemmas, a dataset of 1,834 dilemma questions based on Indonesia's Pancasila values, to evaluate the value alignment of large language models (LLMs). Testing 50 LLMs, the study found that all models scored below 0.5 on Probability Match Score and struggled most with dilemmas related to Religion and Unity, indicating a significant gap in their ability to capture Indonesian human values.

Why it matters: This work highlights the importance of culturally-specific benchmarks for evaluating the value alignment of AI systems, especially in non-Western contexts.

Full story at: arXiv Computation and Language