English

AQAScore: Evaluating Semantic Alignment in Text-to-Audio Generation via Audio Question Answering

Audio and Speech Processing 2026-01-22 v1 Artificial Intelligence Computation and Language Machine Learning Sound

Abstract

Although text-to-audio generation has made remarkable progress in realism and diversity, the development of evaluation metrics has not kept pace. Widely-adopted approaches, typically based on embedding similarity like CLAPScore, effectively measure general relevance but remain limited in fine-grained semantic alignment and compositional reasoning. To address this, we introduce AQAScore, a backbone-agnostic evaluation framework that leverages the reasoning capabilities of audio-aware large language models (ALLMs). AQAScore reformulates assessment as a probabilistic semantic verification task; rather than relying on open-ended text generation, it estimates alignment by computing the exact log-probability of a "Yes" answer to targeted semantic queries. We evaluate AQAScore across multiple benchmarks, including human-rated relevance, pairwise comparison, and compositional reasoning tasks. Experimental results show that AQAScore consistently achieves higher correlation with human judgments than similarity-based metrics and generative prompting baselines, showing its effectiveness in capturing subtle semantic inconsistencies and scaling with the capability of underlying ALLMs.

Keywords

Cite

@article{arxiv.2601.14728,
  title  = {AQAScore: Evaluating Semantic Alignment in Text-to-Audio Generation via Audio Question Answering},
  author = {Chun-Yi Kuan and Kai-Wei Chang and Hung-yi Lee},
  journal= {arXiv preprint arXiv:2601.14728},
  year   = {2026}
}

Comments

Manuscript in progress

R2 v1 2026-07-01T09:13:38.795Z