English

SongBench: A Fine-Grained Multi-Aspect Benchmark for Song Quality Assessment

Audio and Speech Processing 2026-04-30 v1 Artificial Intelligence Sound

Abstract

Recent advancements in Text-to-Song generation have enabled realistic musical content production, yet existing evaluation benchmarks lack the professional granularity to capture multi-dimensional aesthetic nuances. In this paper, we propose SongBench, a specialized framework for fine-grained song assessment across seven key dimensions: Vocal, Instrument, Melody, Structure, Arrangement, Mixing, and Musicality. Utilizing this framework, we construct an expert-annotated database comprising 11,717 samples from state-of-the-art models, labeled by music professionals. Extensive experimental results demonstrate that SongBench achieves high correlation with expert ratings. By revealing fine-grained performance gaps in current state-of-the-art models, SongBench serves as a diagnostic benchmark to steer the development toward more professional and musically coherent song generation.

Keywords

Cite

@article{arxiv.2604.25937,
  title  = {SongBench: A Fine-Grained Multi-Aspect Benchmark for Song Quality Assessment},
  author = {Dapeng Wu and Shun Lei and Wei Tan and Guangzheng Li and Yunzhe Wang and Huaicheng Zhang and Lishi Zuo and Zhiyong Wu},
  journal= {arXiv preprint arXiv:2604.25937},
  year   = {2026}
}