Cultural reward model helps AI behave appropriately worldwide
CRISP: Cultural Reward Modeling for Implicit Situated Propriety
Computation and Language
Summary
As AI language models are used globally, they need to understand different cultural rules to act properly in each place. The authors developed a system called CRISP-RM that helps AI judge whether actions fit local norms in open social situations. They also added a method called Norm Grounding Supervision to better teach the AI what matters in each culture. Their experiments show this approach works better than previous general methods at making AI responses culturally appropriate beyond just polite or fluent language.
What this means in practice
- •For chatbot developers: Improve chatbots’ ability to respond appropriately across different cultural backgrounds in open conversational settings.
- •For international customer support teams: Enhance automated response systems to align better with customers’ cultural expectations and avoid misunderstandings.
Authors
Zekun Yuan, Yangfan Ye, Baohang Li, Shuaibo Zhao, Zekun Zhou, Ziming Li, Qichen Hong, Kun Chen, Xiaocheng Feng
Abstract
As large language models (LLMs) are increasingly deployed across countries and regions, the ability to recognize and respond appropriately to diverse cultural contexts becomes increasingly important. However, existing research has largely focused on cultural knowledge or tasks with predefined response spaces, while open-ended culturally situated behavior remains comparatively underexplored. In this work, we introduce CRISP-RM, a culturally situated reward model that assigns rewards according to cultural appropriateness in open-ended social scenarios. During policy optimization, we further introduce Norm Grounding Supervision (NGS), providing guidance that enhances the policy's sensitivity to relevant cultural norms. To construct culturally situated data, we employ a collaborative multi-agent framework that instantiates implicit cultural norms into diverse social scenarios and further curate NormCompass as a dedicated testbed. We conduct comprehensive experiments to evaluate the effectiveness of CRISP-RM in both reward modeling and policy optimization. Best-of-\(N\) experiments show that CRISP-RM consistently outperforms strong general reward models. During GRPO policy optimization, CRISP-RM generally improves culturally situated behavior, while incorporating NGS yields further gains. Further analyses demonstrate the advantages of CRISP-RM in distinguishing culturally appropriate behavior beyond superficial fluency and politeness, while NGS provides complementary gains during policy optimization by improving norm grounding.