中文

人类与大型音频语言模型在社会语言处理之间的差异:来自模型-大脑对齐的证据

计算与语言 2025-10-28 v2 神经元与认知

摘要

语音AI开发面临独特挑战,需处理语言和非语言信息。本研究比较了大型音频语言模型(LALMs)和人类在语音理解中整合说话人特征的能力,探讨LALMs是否以与人类认知机制相似的方式处理说话人上下文化语言。我们将两种LALMs(Qwen2-Audio和Ultravox 0.5)的处理模式与人类EEG响应进行比较。利用模型的surprisal和entropy指标,我们分析了这些模型对说话人-内容不一致内容的敏感性,包括社会刻板印象违反(例如,男性声称常规做男性指甲)和生物知识违反(例如,男性声称怀孕)。结果显示,Qwen2-Audio对说话人不一致内容表现出增加的surprisal,其surprisal值显著预测了人类的N400反应,而Ultravox 0.5对说话人特征的敏感性有限。重要的是,两种模型都未复制人类对社会违反(引发N400效应)和生物违反(引发P600效应)的处理区别。这些发现揭示了当前LALMs在处理说话人上下文化语言方面的潜力与局限,并表明人类与LALMs之间存在社会语言处理机制的差异。

关键词

引用

@article{arxiv.2503.19586,
  title  = {Distinct social-linguistic processing between humans and large audio-language models: Evidence from model-brain alignment},
  author = {Hanlin Wu and Xufeng Duan and Zhenguang Cai},
  journal= {arXiv preprint arXiv:2503.19586},
  year   = {2025}
}

备注

Hanlin Wu, Xufeng Duan, and Zhenguang Cai. 2025. Distinct social-linguistic processing between humans and large audio-language models: Evidence from model-brain alignment. In Proceedings of the Workshop on Cognitive Modeling and Computational Linguistics, pages 135-143, Albuquerque, New Mexico, USA. Association for Computational Linguistics. https://aclanthology.org/2025.cmcl-1.18/