English

MusiCRS: Benchmarking Audio-Centric Conversational Recommendation

Sound 2026-01-26 v2 Multimedia

Abstract

Conversational recommendation has advanced rapidly with large language models (LLMs), yet music remains a uniquely challenging domain in which effective recommendations require reasoning over audio content beyond what text or metadata can capture. We present MusiCRS, the first benchmark for audio-centric conversational recommendation that links authentic user conversations from Reddit with corresponding tracks. MusiCRS includes 477 high-quality conversations spanning diverse genres (classical, hip-hop, electronic, metal, pop, indie, jazz), with 3,589 unique musical entities and audio grounding via YouTube links. MusiCRS supports evaluation under three input modality configurations: audio-only, query-only, and audio+query, allowing systematic comparison of audio-LLMs, retrieval models, and traditional approaches. Our experiments reveal that current systems struggle with cross-modal integration, with optimal performance frequently occurring in single-modality settings rather than multimodal configurations. This highlights fundamental limitations in cross-modal knowledge integration, as models excel at dialogue semantics but struggle when grounding abstract musical concepts in audio. To facilitate progress, we release the MusiCRS dataset (https://huggingface.co/datasets/rohan2810/MusiCRS), evaluation code (https://github.com/rohan2810/musiCRS), and comprehensive baselines.

Keywords

Cite

@article{arxiv.2509.19469,
  title  = {MusiCRS: Benchmarking Audio-Centric Conversational Recommendation},
  author = {Rohan Surana and Amit Namburi and Gagan Mundada and Abhay Lal and Zachary Novack and Julian McAuley and Junda Wu},
  journal= {arXiv preprint arXiv:2509.19469},
  year   = {2026}
}

Comments

5 pages

R2 v1 2026-07-01T05:52:56.769Z