中文

推理塑造对齐:基于文化规范探讨大规模推理模型的文化对齐

人工智能 2025-11-18 v1

摘要

大规模推理模型的先进推理能力使其能够通过深思熟虑的过程全面理解并应用安全政策,从而提高模型的安全性。除了安全性之外,这些模型还必须能够反映人类在各种文化中所体现的多样化价值观。本文提出了基于文化规范的文化对齐框架(Cultural Norm-based Cultural Alignment, CNCA),该框架使模型能够利用其强大的推理能力与文化规范保持一致。具体而言,我们提出了三种方法,从有限的调查数据中自动挖掘文化规范,并探讨了有效利用这些规范以提高文化对齐的方式。 examined two alignment paradigms: an in-context alignment method, where cultural norms are explicitly integrated into the user context, and a fine-tuning-based method, which internalizes norms through enhanced Chain-of-Thought training data. Comprehensive experiments demonstrate the effectiveness of these methods, highlighting that models with stronger reasoning capabilities benefit more from cultural norm mining and utilization. Our findings emphasize the potential for reasoning models to better reflect diverse human values through culturally informed alignment strategies.

关键词

引用

@article{arxiv.2511.13359,
  title  = {Reasoning Shapes Alignment: Investigating Cultural Alignment in Large Reasoning Models with Cultural Norms},
  author = {Yuhang Wang and Yanxu Zhu and Jitao Sang},
  journal= {arXiv preprint arXiv:2511.13359},
  year   = {2025}
}