Whisper模型在儿童语音识别中的自适应
音频与语音处理
2023-07-26 v1 人工智能
摘要
自动语音识别(ASR)系统往往难以转录儿童语音,因为缺乏准确训练儿童友好型ASR模型所需的大规模儿童语音数据集。然而,存在大量已标注的成人语音数据集,这些数据集曾被用于创建多语言ASR模型,例如Whisper。我们的工作旨在探索此类模型是否可适配于儿童语音,以改善面向儿童的ASR性能。此外,我们将Whisper的儿童语音自适应与微调后的自监督模型(如wav2vec2)进行比较。我们表明,与未微调的Whisper模型相比,在儿童语音上微调Whisper可显著提升其在儿童语音上的ASR性能。另外,利用已在儿童语音上微调的自监督Wav2vec2模型优于Whisper微调。
引用
@article{arxiv.2307.13008,
title = {Adaptation of Whisper models to child speech recognition},
author = {Rishabh Jain and Andrei Barcovschi and Mariam Yiwere and Peter Corcoran and Horia Cucu},
journal= {arXiv preprint arXiv:2307.13008},
year = {2023}
}
备注
Accepted in Interspeech 2023