中文

论语义对齐语音表示在口语理解中的应用

计算与语言 2022-10-12 v1 声音 音频与语音处理

摘要

本文研究语义对齐的语音表示在端到端口语理解(SLU)中的应用。我们采用最近提出的 SAMU-XLSR 模型,该模型旨在生成在话语层面捕获语义的单一嵌入,并在不同语言间实现语义对齐。该模型将声学帧级语音表示学习模型(XLS-R)与语言无关 BERT 句子嵌入(LaBSE)模型相结合。我们表明,使用 SAMU-XLSR 模型代替初始 XLS-R 模型可显著提升端到端 SLU 框架中的性能。最后,我们展示了该模型在 SLU 语言可移植性方面的优势。

关键词

引用

@article{arxiv.2210.05291,
  title  = {On the Use of Semantically-Aligned Speech Representations for Spoken Language Understanding},
  author = {Gaëlle Laperrière and Valentin Pelloin and Mickaël Rouvier and Themos Stafylakis and Yannick Estève},
  journal= {arXiv preprint arXiv:2210.05291},
  year   = {2022}
}

备注

Accepted in IEEE SLT 2022. This work was performed using HPC resources from GENCI/IDRIS (grant 2022 AD011012565) and received funding from the EU H2020 research and innovation programme under the Marie Sklodowska-Curie ESPERANTO project (grant agreement No 101007666), through the SELMA project (grant No 957017) and from the French ANR through the AISSPER project (ANR-19-CE23-0004)