基于重合成时长建模的真实场景语音情感转换增强
音频与语音处理
2025-08-18 v1
摘要
语音情感转换旨在修改输入语音所表达的情感,同时保留词汇内容和说话人身份。近期,生成式建模方法在改变基频、频谱包络和能量等局部声学特性方面取得了有前景的结果,但往往缺乏对声音时长的控制能力。为解决这一问题,我们提出了一种基于重合成的离散内容表征的时长建模框架,能够在无需平行数据的情况下修改语音时长以反映目标情感,并实现可控的语音速率。实验结果表明,所提出的时长建模框架的加入显著增强了真实场景MSP-Podcast数据集中的情感表现力。分析表明,低唤醒度情感与更长的时长和更慢的语音速率相关,而高唤醒度情感则产生更短、更快的语音。
引用
@article{arxiv.2508.11535,
title = {Enhancing In-the-Wild Speech Emotion Conversion with Resynthesis-based Duration Modeling},
author = {Navin Raj Prabhu and Danilo de Oliveira and Nale Lehmann-Willenbrock and Timo Gerkmann},
journal= {arXiv preprint arXiv:2508.11535},
year = {2025}
}
备注
Copyright 2025 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works