中文

论使用预训练嵌入进行音频分类时最优时间支撑的选择

声音 2023-12-22 v1 人工智能 音频与语音处理

摘要

当前 SOTA 音频分析系统依赖于预训练嵌入模型,通常将其作为(冻结的)特征提取器即插即用。为一系列任务选择最佳模型是近期许多出版物的主题。然而,这些工作中常被忽视的一个方面是提取嵌入所考虑的音频输入时长的影响,我们将其称为时间支撑(TS)。在本工作中,我们研究了 TS 对成熟或新兴预训练嵌入的影响,这些嵌入被选择以代表不同类型的架构和学习范式。我们使用乐器和环境声音数据集进行评估,即 OpenMIC、TAU Urban Acoustic Scenes 2020 Mobile 和 ESC-50。我们特别强调,基于 Audio Spectrogram Transformer 的系统(PaSST 和 BEATs)在较小 TS 下依然有效,从而允许大幅降低内存和计算成本。此外,我们表明通过选择最优 TS,我们在所有任务上均取得了有竞争力的结果。特别是,我们在 OpenMIC 上使用 BEATs 和 PaSST 在无任何微调的情况下改进了 SOTA 结果。

关键词

引用

@article{arxiv.2312.14005,
  title  = {On the choice of the optimal temporal support for audio classification with Pre-trained embeddings},
  author = {Aurian Quelennec and Michel Olvera and Geoffroy Peeters and Slim Essid},
  journal= {arXiv preprint arXiv:2312.14005},
  year   = {2023}
}

备注

Copyright 2024 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works