English

A Benchmark for Audio Reasoning Capabilities of Multimodal Large Language Models

Sound 2026-01-28 v1 Artificial Intelligence

Abstract

The present benchmarks for testing the audio modality of multimodal large language models concentrate on testing various audio tasks such as speaker diarization or gender identification in isolation. Whether a multimodal model can answer the questions that require reasoning skills to combine audio tasks of different categories, cannot be verified with their use. To address this issue, we propose Audio Reasoning Tasks (ART), a new benchmark for assessing the ability of multimodal models to solve problems that require reasoning over audio signal.

Keywords

Cite

@article{arxiv.2601.19673,
  title  = {A Benchmark for Audio Reasoning Capabilities of Multimodal Large Language Models},
  author = {Iwona Christop and Mateusz Czyżnikiewicz and Paweł Skórzewski and Łukasz Bondaruk and Jakub Kubiak and Marcin Lewandowski and Marek Kubis},
  journal= {arXiv preprint arXiv:2601.19673},
  year   = {2026}
}

Comments

31 pages, 2 figures, accepted to EACL 2026

R2 v1 2026-07-01T09:22:24.497Z