What Changed
Traditional approaches to equipping language models (LMs) with moral reasoning capabilities often fall short in multilingual and multicultural settings. Existing evaluation benchmarks frequently rely on direct translation, failing to capture culture-specific moral nuances. Inference-time methods are typically English-centric and lack grounding in established moral theories. Furthermore, training these models often necessitates expensive supervision. The new research introduces a three-pronged solution to these challenges:
MCLASH Benchmark: A new multilingual moral decision-making benchmark designed to assess culturally situated moral intuitions and social norms across various languages. Unlike previous benchmarks, MCLASH aims to include culture-specific items rather than relying solely on direct translations.
MET (Multilingual Ethics with Theory-grounded reasoning): A two-step prompting methodology. This method leverages expert-curated, theory-based grounds derived from psychology and philosophy. In the first step, the model selects situation- and culture-specific moral grounds. In the second step, it reasons over these selected grounds in the user's native language.
MET-D (MET-Distillation): An enhancement to the MET framework's second reasoning step. MET-D utilizes a self-distillation training stage, eliminating the need for external supervision from stronger models or human annotators, thereby reducing training costs and complexity.






