Experts create alignment data for language models on Islamic ethics
From Normative Frameworks to Alignment Data: Constructing and Evaluating SFT and Preference Data
Computation and LanguageArtificial Intelligence
Summary
Aligning large language models with specific ethical principles requires turning abstract ideas into examples that models can learn from. The authors worked with experts in Islamic ethics to generate guided examples and preference comparisons in Arabic and English over a year. They tested models trained on this data and found improvements in following the expert-aligned principles without losing general language capabilities. This shows it is possible to systematically use expert knowledge to guide model behavior in certain ethical frameworks.
What this means in practice
- •For ai ethics teams: Design custom language models that follow specific ethical and cultural guidelines using expert-curated training examples and preferences.
- •For multilingual ai developers: Create Arabic-English language models better aligned with culturally relevant values by incorporating expert-driven alignment data.
Authors
Husrev Taha Sencar, Rezart Beka, Danish Naeem, Seda Ozalkan, Majd Hawasly, Ji Lucas, Ala AlFuqaha, Mohamed Abdallah, Recep Senturk
Abstract
Aligning language models with a specified normative framework requires translating abstract principles into concrete examples and preference signals from which models can learn. We present an expert-driven methodology for constructing such alignment data and apply it to a normative framework grounded in Islamic ethical, theological, and jurisprudential traditions. Over approximately one year, seven domain experts systematically probed language models to identify alignment deficiencies, curated desired responses, and constructed preference pairs from model outputs and expert judgments. The resulting Arabic-English datasets contain approximately 2.8K supervised fine-tuning (SFT) examples and 5.4K preference pairs spanning a broad range of normative domains. We evaluate the datasets through controlled post-training experiments comparing a Baseline model with models incorporating the curated SFT data alone and both the SFT and preference data. In blind expert evaluation on 150 separately constructed prompts, the model trained with the curated SFT data was preferred over the Baseline in 51.3% of assessor judgments, compared with 14.4% in the opposite direction (p < .001 at the prompt level). Adding the preference data resulted in a smaller difference, with the model trained with both datasets preferred over the SFT model in 28.0% of judgments versus 20.9% in the opposite direction; this difference was not statistically significant at the prompt level (p = .166). Standard Arabic and English benchmarks show no broad degradation in general-purpose capabilities. These results demonstrate how expert-defined normative principles can be systematically operationalized into alignment data and evaluated through controlled model training.