中文

探索用于情感发声预测少样本个性化的说话人注册

声音 2022-06-22 v2 机器学习 音频与语音处理

摘要

本文中,我们探索一种用于情感发声预测的少样本个性化新架构。核心贡献是一个“注册”编码器,其利用目标说话人的两个无标签样本来调整情感编码器的输出;该调整基于点积注意力,从而有效地充当一种“软”特征选择形式。情感与注册编码器基于两种标准音频架构:CNN14与CNN10。这两个编码器进一步被引导去遗忘或学习辅助情感及/或说话人信息。我们的最佳方法在ExVo Few-Shot开发集上取得了.650.650的CCC,较我们基线CNN14的.634.634 CCC提高了2.5%2.5\%

关键词

引用

@article{arxiv.2206.06680,
  title  = {Exploring speaker enrolment for few-shot personalisation in emotional vocalisation prediction},
  author = {Andreas Triantafyllopoulos and Meishu Song and Zijiang Yang and Xin Jing and Björn W. Schuller},
  journal= {arXiv preprint arXiv:2206.06680},
  year   = {2022}
}

备注

Proceedings of the ICML Expressive Vocalizations Workshop and Competition held in conjunction with the $\mathit{39}^{th}$ International Conference on Machine Learning, Copyright 2022 by the author(s)