persoDA:面向个性化语音识别的个人化数据增强
摘要
数据增强(DA)在自动语音识别(ASR)模型的训练中普遍使用。DA 提供了增加数据可变性、鲁棒性和对不同声学扰动的泛化能力。最近的研究表明,针对移动设备上的个性化 ASR 模型可提高词汇错误率(WER)。本文评估了在此背景下的数据增强,并提出了 persoDA;一种由用户数据驱动的 DA 方法,用于个性化 ASR。persoDA 旨�通过针对最终用户声学特征进行优化来增强训练数据,与基于多条件训练(MCT)的标准增强方法相比,后者应用随机混响和噪声。我们使用在 Librispeech 上训练并针对 VOICES 个性化的 conformer 基线进行评估,表明 persoDA 相比使用标准数据增强(使用随机噪声和混响)可实现 13.9% 的相对 WER 降低。此外,persoDA 显示比 MCT 快 16% 到 20% 的收敛速度。
引用
@article{arxiv.2501.09113,
title = {persoDA: Personalized Data Augmentation for Personalized ASR},
author = {Pablo Peso Parada and Spyros Fontalis and Md Asif Jalal and Karthikeyan Saravanan and Anastasios Drosou and Mete Ozay and Gil Ho Lee and Jungin Lee and Seokyeong Jung},
journal= {arXiv preprint arXiv:2501.09113},
year = {2025}
}
备注
ICASSP'25-Copyright 2025 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works