中文

SPGISpeech 2.0:用于说话人标记转录的金融领域多说话人音频

声音 2025-08-08 v1 计算与语言 音频与语音处理

摘要

我们介绍 SPGISpeech 2.0,这是一个适用于说话人标记转录的数据集。SPGISpeech 2.0 在保持原始 SPGISpeech 数据集核心特征的同时,提高了适用建模任务的多样性:音频片段及其对应的完整格式文本转录,可用于端到端自动语音识别 (ASR)。SPGISpeech 2.0 包含 3,780 小时专业转录的财报通话记录。此外,数据集包含每个音频片段的通话信息和说话人信息,便于多说话人 ASR。我们通过在 SPGISpeech 2.0 上微调热门语音识别模型,验证了 SPGISpeech 2.0 在说话人标记 ASR 性能方面的提升。该数据集以非商业用途免费发布,我们期望 SPGISpeech 2.0 能推动语音识别技术的进步,并激发广泛的研究应用。

关键词

引用

@article{arxiv.2508.05554,
  title  = {SPGISpeech 2.0: Transcribed multi-speaker financial audio for speaker-tagged transcription},
  author = {Raymond Grossman and Taejin Park and Kunal Dhawan and Andrew Titus and Sophia Zhi and Yulia Shchadilova and Weiqing Wang and Jagadeesh Balam and Boris Ginsburg},
  journal= {arXiv preprint arXiv:2508.05554},
  year   = {2025}
}

备注

To be presented at Interspeech 2025