探索用于情感发声预测少样本个性化的说话人注册
声音
2022-06-22 v2 机器学习
音频与语音处理
摘要
本文中,我们探索一种用于情感发声预测的少样本个性化新架构。核心贡献是一个“注册”编码器,其利用目标说话人的两个无标签样本来调整情感编码器的输出;该调整基于点积注意力,从而有效地充当一种“软”特征选择形式。情感与注册编码器基于两种标准音频架构:CNN14与CNN10。这两个编码器进一步被引导去遗忘或学习辅助情感及/或说话人信息。我们的最佳方法在ExVo Few-Shot开发集上取得了的CCC,较我们基线CNN14的 CCC提高了。
引用
@article{arxiv.2206.06680,
title = {Exploring speaker enrolment for few-shot personalisation in emotional vocalisation prediction},
author = {Andreas Triantafyllopoulos and Meishu Song and Zijiang Yang and Xin Jing and Björn W. Schuller},
journal= {arXiv preprint arXiv:2206.06680},
year = {2022}
}
备注
Proceedings of the ICML Expressive Vocalizations Workshop and Competition held in conjunction with the $\mathit{39}^{th}$ International Conference on Machine Learning, Copyright 2022 by the author(s)