中文

Optimizing Speech Multi-View Feature Fusion through Conditional Computation

音频与语音处理 2025-01-15 v1 人工智能 计算与语言 声音

摘要

近期的进展突显了自监督学习 (SSL) 特征在各种语音相关任务中的有效性,提供了轻量且通用的多视图语音表示。然而,我们的研究发现,尽管 SSL 特征加快了模型收敛,但它们与诸如 FBanks 之类的传统频谱特征在更新方向上存在冲突。作为响应,我们提出了一种基于条件计算的新型通用特征融合框架,特点是梯度敏感性门控网络和多阶段 dropout 策略。该框架缓解了特征冲突,增强了模型对多视图输入特征的鲁棒性。通过整合 SSL 和频谱特征,我们的方法加快了收敛速度,并在 MUSTC 数据集上的多个语音翻译任务中保持与频谱模型相当的性能。

关键词

引用

@article{arxiv.2501.08057,
  title  = {Optimizing Speech Multi-View Feature Fusion through Conditional Computation},
  author = {Weiqiao Shan and Yuhao Zhang and Yuchen Han and Bei Li and Xiaofeng Zhao and Yuang Li and Min Zhang and Hao Yang and Tong Xiao and Jingbo Zhu},
  journal= {arXiv preprint arXiv:2501.08057},
  year   = {2025}
}

备注

ICASSP 2025