Optimizing Speech Multi-View Feature Fusion through Conditional Computation
音频与语音处理
2025-01-15 v1 人工智能
计算与语言
声音
摘要
近期的进展突显了自监督学习 (SSL) 特征在各种语音相关任务中的有效性,提供了轻量且通用的多视图语音表示。然而,我们的研究发现,尽管 SSL 特征加快了模型收敛,但它们与诸如 FBanks 之类的传统频谱特征在更新方向上存在冲突。作为响应,我们提出了一种基于条件计算的新型通用特征融合框架,特点是梯度敏感性门控网络和多阶段 dropout 策略。该框架缓解了特征冲突,增强了模型对多视图输入特征的鲁棒性。通过整合 SSL 和频谱特征,我们的方法加快了收敛速度,并在 MUSTC 数据集上的多个语音翻译任务中保持与频谱模型相当的性能。
引用
@article{arxiv.2501.08057,
title = {Optimizing Speech Multi-View Feature Fusion through Conditional Computation},
author = {Weiqiao Shan and Yuhao Zhang and Yuchen Han and Bei Li and Xiaofeng Zhao and Yuang Li and Min Zhang and Hao Yang and Tong Xiao and Jingbo Zhu},
journal= {arXiv preprint arXiv:2501.08057},
year = {2025}
}
备注
ICASSP 2025