中文

如何连接语音基础模型与大型语言模型?哪些因素重要,哪些因素不重要

计算与语言 2025-06-04 v3 人工智能 机器学习

摘要

大型语言模型 (LLM) 取得的卓越性能正在推动研究人员致力于将其应用于各种任务和输入模态。在语音到文本 (S2T) 任务中,新兴的解决方案是通过适配器模块将语音基础模型 (SFM) 的编码器输出投影到 LLM 嵌入空间。然而,尚无工作调查下游任务性能如何依赖于各个组件 (SFM、适配器、LLM),以及最佳适配器设计是否取决于所选择的 SFM 和 LLM。为填补这一空白,我们评估了 5 个适配器模块、2 个 LLM (Mistral 和 Llama) 和 2 个 SFM (Whisper 和 SeamlessM4T) 在两种广泛的 S2T 任务(自动语音识别和语音翻译)上的组合。我们的结果表明,SFM 在下游性能中发挥关键作用,而适配器选择的影响相对适中,且取决于 SFM 和 LLM。

关键词

引用

@article{arxiv.2409.17044,
  title  = {How to Connect Speech Foundation Models and Large Language Models? What Matters and What Does Not},
  author = {Francesco Verdini and Pierfrancesco Melucci and Stefano Perna and Francesco Cariaggi and Marco Gaido and Sara Papi and Szymon Mazurek and Marek Kasztelnik and Luisa Bentivogli and Sébastien Bratières and Paolo Merialdo and Simone Scardapane},
  journal= {arXiv preprint arXiv:2409.17044},
  year   = {2025}
}

备注

Submitted to Interspeech 2025