音乐音视频问答需要专门的多模态设计
声音
2026-04-13 v2 计算机视觉与模式识别
多媒体
音频与语音处理
摘要
虽然最近的多模态大语言模型在 general multimodal tasks 上表现突出,但 specialized 领域如音乐却需要量身打造的方法。音乐音视频问答 (Music AVQA) 尤其突显了这一点,具有连续、密集编排的音视频内容、复杂的时间动态以及关键的 domain-specific 知识需求。通过对 Music AVQA 数据集和方法的系统分析,本文确定 specialized 输入处理、融合专门空间时间设计的架构以及 music-specific 建模策略是该领域成功的关键。我们的研究通过揭示与强性能相关的有效设计模式,提出具体的未来方向以整合音乐先验,并旨在为推进多模态音乐理解奠定稳健的基础。我们旨在鼓励此领域的进一步研究,提供一个 GitHub 仓库收录相关作品:https://github.com/WenhaoYou1/Survey4MusicAVQA。
引用
@article{arxiv.2505.20638,
title = {Music Audio-Visual Question Answering Requires Specialized Multimodal Designs},
author = {Wenhao You and Xingjian Diao and Wenjun Huang and Chunhui Zhang and Keyi Kong and Weiyi Wu and Chiyu Ma and Zhongyu Ouyang and Tingxuan Wu and Ming Cheng and Soroush Vosoughi and Jiang Gui},
journal= {arXiv preprint arXiv:2505.20638},
year = {2026}
}
备注
Accepted to Annual Meeting of the Association for Computational Linguistics (ACL 2026). The first two authors contributed equally