人类在交互时听什么?基于人类选择性听力实验的语音识别系统评估
计算与语言
2025-10-09 v2
摘要
spoken dialogue systems (SDSs) 采用自动语音识别 (ASR) 作为其管道的前端。ASR 在 SDSs 中的作用是识别与响应生成相关的用户语音中的信息。考察人类的选择性听力,这指的是在说话过程中专注于并倾听对话中重要部分的能力,将有助于我们确定 SDSs 所需的 ASR 能力并进行评估。在本研究中,我们通过比较用于生成对话响应的人类转录与参考转录,实验性地确认了人类在生成对话响应时进行选择性听力。基于我们的实验结果,我们讨论了一种新的 ASR 评估方法,该方法利用人类选择性听力,可识别 ASR 系统与人类之间转录能力的差距。
关键词
引用
@article{arxiv.2508.04402,
title = {What Do Humans Hear When Interacting? Experiments on Selective Listening for Evaluating ASR of Spoken Dialogue Systems},
author = {Kiyotada Mori and Seiya Kawano and Chaoran Liu and Carlos Toshinori Ishi and Angel Fernando Garcia Contreras and Koichiro Yoshino},
journal= {arXiv preprint arXiv:2508.04402},
year = {2025}
}
备注
Revised version with Table 5 updated for ADP, NUM, PROPN, and PRON