改进ECAPA-TDNN多尺度特征融合与注意力增强的婴儿啼叫情感识别
音频与语音处理
2025-06-24 v1
摘要
婴儿啼叫情感识别对于育儿与医疗应用至关重要,但面临诡异情感差异、噪声干扰和数据有限等挑战。现有方法缺乏有效整合多尺度特征与时频关系的能力。本研究提出一种改进版强调通道注意力、传播与聚合于时延神经网络(ECAPA-TDNN)的方法,集成多尺度特征融合与注意力增强。实验在公开数据集上表明,所提方法实现了82.20%的准确率,参数数量为1.43 MB,FLOPs为0.32 Giga。此外,我们的方法在准确率方面优于基线方法。代码地址为 https://github.com/kkpretend/IETMA。
引用
@article{arxiv.2506.18402,
title = {Infant Cry Emotion Recognition Using Improved ECAPA-TDNN with Multiscale Feature Fusion and Attention Enhancement},
author = {Junyu Zhou and Yanxiong Li and Haolin Yu},
journal= {arXiv preprint arXiv:2506.18402},
year = {2025}
}
备注
Accepted for publication on Interspeech 2025. 5 pages, 2 tables and 7 figures