Pancasila-Dilemmas: Evaluating Large Language Models on Indonesian Human Value Dilemmas Grounded in Pancasila

2026-07-20Computation and Language

Computation and Language
AI summary

The authors created a new dataset called Pancasila-Dilemmas with 1,834 questions based on Indonesian values to test how well large language models (LLMs) align with local cultural beliefs. They focused on dilemmas related to five key Pancasila values: Religion, Humanity, Unity, Democracy, and Social Justice. After testing 50 different LLMs, the authors found that models struggled especially with dilemmas about Religion and Unity, showing they don’t capture Indonesian values well. The dataset and evaluation provide a way to better understand and improve how LLMs handle culturally specific values.

Large Language ModelsValue AlignmentPancasilaIndonesian ValuesEthics in AICultural DilemmasDataset EvaluationProbability Match ScoreMax-Vote Agreement Score
Authors
Supryadi, Irfan, Julianti, Darren Keanly Martin, Jayvin Fernando, Yuqi Ren, Deyi Xiong
Abstract
The value alignment of large language models (LLMs) is crucial for ensuring responses align with human intention and value preferences. However, most evaluations of value alignment focus on Western or universal values, while assessments grounded in the value systems of specific countries remain scarce. In this paper, we introduce Pancasila-Dilemmas, an evaluation dataset of 1,834 questions derived from Indonesian news, classified by 5 values of Pancasila: Religion, Humanity, Unity, Democracy, and Social Justice. This dataset reflects daily life in Indonesia, making it suitable for measuring the value alignment of LLMs deployed for Indonesia. To ensure a more rigorous evaluation, we choose scenarios containing dilemmas. The dataset is proofread by native speakers and answered by 5 diverse Indonesian citizens. We evaluate 50 closed- and open-source LLMs on our dataset. Results reveal that all evaluated LLMs achieves less than 0.5 Probability Match Score (PMS) and 0.72 Max-Vote Agreement Score (MVAS). Compared by each values, LLMs mostly struggle in Religion and Unity dilemma cases. This highlights a significant gap in capturing Indonesian values. The dataset is publicly available at https://github.com/tjunlp-lab/Pancasila-Dilemmas.