Speech-MASSIVE:面向SLU及其扩展任务的多语言语音数据集
计算与语言
2024-08-08 v1 声音
音频与语音处理
摘要
我们提出 Speech-MASSIVE,一个多语言口语语言理解(Spoken Language Understanding, SLU)数据集,包含 MASSIVE 文本语料库部分的语音版本。Speech-MASSIVE 覆盖来自不同语言家族的 12 种语言,继承 MASSIVE 用于意图预测和槽填充任务的注释。我们的扩展由大规模多语言 SLU 数据集的稀缺性以及评估基础模型(大语言模型、语音编码器)在不同语言和任务上的通用性推动。我们提供一个多模态、多任务、多语言数据集,并在各种训练场景(零样本、少样本和完全微调)中使用级联架构和端到端架构报告 SLU 基线。此外,我们展示了 Speech-MASSIVE 适用于基准测试其他任务,如语音转录、语言识别和语音翻译。数据集、模型和代码均公开于:https://github.com/hlt-mt/Speech-MASSIVE
引用
@article{arxiv.2408.03900,
title = {Speech-MASSIVE: A Multilingual Speech Dataset for SLU and Beyond},
author = {Beomseok Lee and Ioan Calapodescu and Marco Gaido and Matteo Negri and Laurent Besacier},
journal= {arXiv preprint arXiv:2408.03900},
year = {2024}
}
备注
Accepted at INTERSPEECH 2024. This version includes the same content but with additional appendices