Papers for

international customer support teams

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Cultural reward model helps AI behave appropriately worldwide

CRISP: Cultural Reward Modeling for Implicit Situated Propriety

Abstract: As large language models (LLMs) are increasingly deployed across countries and regions, the ability to recognize and respond appropriately to diverse cultural contexts becomes increasingly important. However, existing research has largely focused on cultural knowledge or tasks with predefined response spaces, while open-ended culturally situated behavior remains comparatively underexplored. In this work, we introduce CRISP-RM, a culturally situated reward model that assigns rewards according to cultural appropriateness in open-ended social scenarios. During policy optimization, we further introduce Norm Grounding Supervision (NGS), providing guidance that enhances the policy's sensitivity to relevant cultural norms. To construct culturally situated data, we employ a collaborative multi-agent framework that instantiates implicit cultural norms into diverse social scenarios and further curate NormCompass as a dedicated testbed. We conduct comprehensive experiments to evaluate the effectiveness of CRISP-RM in both reward modeling and policy optimization. Best-of-\(N\) experiments show that CRISP-RM consistently outperforms strong general reward models. During GRPO policy optimization, CRISP-RM generally improves culturally situated behavior, while incorporating NGS yields further gains. Further analyses demonstrate the advantages of CRISP-RM in distinguishing culturally appropriate behavior beyond superficial fluency and politeness, while NGS provides complementary gains during policy optimization by improving norm grounding.

Mon 28 SeptComputation and Language
The gist
As AI language models are used globally, they need to understand different cultural rules to act properly in each place. The authors developed a system called CRISP-RM that helps AI judge whether actions fit local norms in open social situations. They also added a method called Norm Grounding Supervision to better teach the AI what matters in each culture. Their experiments show this approach works better than previous general methods at making AI responses culturally appropriate beyond just polite or fluent language.
Open → 2609.34345v1