中文

六 llamas:通过LoRA适配的语言模型进行比较宗教伦理

人工智能 2026-04-21 v1

摘要

我们提出Six Llamas,对大语言模型在不同宗教语料库上微调后是否编码系统性不同伦理推理模式进行比较研究。构建了六个Meta-Llama-3.1-8B变体:一个未修改的控制模型和五个仅在神圣和神学文本中训练的Christianity、Islam、Judaism、Hinduism或Buddhism的LoRA适配模型。所有六个模型都用相同的17个标准化伦理提示器 probed,涵盖道德困境、博弈论情景、公共政策问题和道德心理学自我评估。为评估鲁棒性和可重复性,我们实施了跨温度采样设计,覆盖十个温度设置。我们计算响应一致性指标、两两模型间一致性率、温度灵敏度系数(跨四个提示域)以及运行间稳定性分析。研究结果表明,LoRA适配模型产生的伦理推理模式(a)与基础模型系统性不同,(b)与其训练传统的道德逻辑一致,(c)在道德哲学空间中沿可解释维度结构化,(d)高共识困境的核心伦理立场在温度变化下保持稳定。拖车问题在所有模型和温度设置下实现100%一致性,而(e)在道德争议域中,传统特定的差异在较高温度下加剧,(f)基础模型表现出最高的整体响应一致性(平均88.3%),表明LoRA适配既引入了传统特定信号,又增加了采样灵敏度。该研究为使用差别训练的语言模型作为文化和伦理分析工具的凝聚比较方法提供了概念验证,并确定了特定的否定性准则和计划扩展。

关键词

引用

@article{arxiv.2604.18404,
  title  = {Six Llamas: Comparative Religious Ethics Through LoRA-Adapted Language Models},
  author = {Chad Coleman and W. Russell Neuman and Manan Shah and Ali Dasdan and Matthew Crispi and Morris Chiang and Zack Leitman and Mustafa Poonawala},
  journal= {arXiv preprint arXiv:2604.18404},
  year   = {2026}
}

备注

51 pages, 14 figures. We present Six Llamas, a comparative study examining whether Llama-3.1-8B models fine-tuned on distinct religious corpora encode systematically different patterns of ethical reasoning. Five LoRA-adapted variants are constructed for Christianity, Islam, Judaism, Hinduism, and Buddhism. For theoretical background on the condensate comparative method, see arXiv:2603.07329