VoiceGuider:通过自导盼提升参数高效说话人自适应文本转语音的跨域性能
声音
2024-12-24 v2 音频与语音处理
摘要
在通过LoRA对说话人自适应文本转语音模型进行参数高效微调时,适应性能可能低于全参数微调的版本,尤其是针对跨域说话人时。本文提出VoiceGuider,一个通过自导盼增强的、参数高效的说话人自适应文本转语音系统,以提升说话人适应性能,缩小与全参数微调模型之间的差距。我们仔细探索了各种加强自导盼的方式,最终找到最佳策略。VoiceGuider因此在极端跨域语音数据上表现出稳健的适应性能。我们提供了演示页面中的可听样本。
引用
@article{arxiv.2409.15759,
title = {VoiceGuider: Enhancing Out-of-Domain Performance in Parameter-Efficient Speaker-Adaptive Text-to-Speech via Autoguidance},
author = {Jiheum Yeom and Heeseung Kim and Jooyoung Choi and Che Hyun Lee and Nohil Park and Sungroh Yoon},
journal= {arXiv preprint arXiv:2409.15759},
year = {2024}
}
备注
IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP), 2025, Demo Page: https://voiceguider.github.io/