中文

AV-SUPERB:面向音视表示模型的多任务评测基准

音频与语音处理 2024-03-20 v2 计算机视觉与模式识别 多媒体 声音

摘要

音视表示学习旨在利用听觉与视觉信息间的相关性,开发具有类人感知的系统。然而,当前模型往往聚焦于有限的一组任务,所学表示的泛化能力尚不明晰。为此,我们提出 AV-SUPERB 基准,可在语音与音频处理中涵盖 5 项音视任务的 7 个数据集上,对单模态音频/视觉及双模态融合表示进行通用评测。我们评估了 5 个近期自监督模型,表明这些模型均无法泛化至所有任务,凸显未来研究改进通用模型性能的必要性。此外,我们展示通过中间任务微调可改进表示,且利用 AudioSet 进行音频事件分类是一项强力的中间任务。我们发布包含评测代码与模型提交平台的基准,以鼓励音视学习的进一步研究。

关键词

引用

@article{arxiv.2309.10787,
  title  = {AV-SUPERB: A Multi-Task Evaluation Benchmark for Audio-Visual Representation Models},
  author = {Yuan Tseng and Layne Berry and Yi-Ting Chen and I-Hsiang Chiu and Hsuan-Hao Lin and Max Liu and Puyuan Peng and Yi-Jen Shih and Hung-Yu Wang and Haibin Wu and Po-Yao Huang and Chun-Mao Lai and Shang-Wen Li and David Harwath and Yu Tsao and Shinji Watanabe and Abdelrahman Mohamed and Chi-Luen Feng and Hung-yi Lee},
  journal= {arXiv preprint arXiv:2309.10787},
  year   = {2024}
}

备注

Accepted to ICASSP 2024; Evaluation Code: https://github.com/roger-tseng/av-superb Submission Platform: https://av.superbbenchmark.org