自监督声学词嵌入的逐层分析:一项关于语音情感识别的研究
计算与语言
2024-02-06 v1 声音
音频与语音处理
摘要
自监督语音模型的有效性已得到验证,然而在不同任务中优化利用其表示仍然具有挑战性。在本研究中,我们深入探讨了声学词嵌入(Acoustic Word Embeddings, AWEs),这是一种源自连续表示的定长特征,以探索其在特定任务中的优势。AWEs 此前已被证明在捕捉声学判别力方面具有效用。鉴于此,我们提议测量 AWEs 与词嵌入之间的逐层相似性,旨在进一步调查 AWEs 内的固有上下文。此外,我们在语音情感识别(Speech Emotion Recognition, SER)的背景下评估了 AWEs 与其他类型语音特征相比的贡献。通过对两个不同语料库 IEMOCAP 和 ESD 的对比实验和逐层准确率分析,我们探讨了 AWEs 与原始自监督表示之间的差异,以及单独使用 AWEs 和结合词嵌入使用 AWEs 的适当方法。我们的发现强调了 AWEs 传达的声学上下文,并展示了通过适当运用 AWEs 可实现极具竞争力的 SER 准确率。
引用
@article{arxiv.2402.02617,
title = {Layer-Wise Analysis of Self-Supervised Acoustic Word Embeddings: A Study on Speech Emotion Recognition},
author = {Alexandra Saliba and Yuanchao Li and Ramon Sanabria and Catherine Lai},
journal= {arXiv preprint arXiv:2402.02617},
year = {2024}
}
备注
Accepted to ICASSP2024 Self-supervision in Audio, Speech and Beyond (SASB) workshop. First two authors contributed equally