English

Hearing Between the Lines: Unlocking the Reasoning Power of LLMs for Speech Evaluation

Computation and Language 2026-01-27 v2

Abstract

Large Language Model (LLM) judges exhibit strong reasoning capabilities but are limited to textual content. This leaves current automatic Speech-to-Speech (S2S) evaluation methods reliant on opaque and expensive Audio Language Models (ALMs). In this work, we propose TRACE (Textual Reasoning over Audio Cues for Evaluation), a novel framework that enables LLM judges to reason over audio cues to achieve cost-efficient and human-aligned S2S evaluation. To demonstrate the strength of the framework, we first introduce a Human Chain-of-Thought (HCoT) annotation protocol to improve the diagnostic capability of existing judge benchmarks by separating evaluation into explicit dimensions: content (C), voice quality (VQ), and paralinguistics (P). Using this data, TRACE constructs a textual blueprint of inexpensive audio signals and prompts an LLM to render dimension-wise judgments, fusing them into an overall rating via a deterministic policy. TRACE achieves higher agreement with human raters than ALMs and transcript-only LLM judges while being significantly more cost-effective. We will release the HCoT annotations and the TRACE framework to enable scalable and human-aligned S2S evaluation.

Keywords

Cite

@article{arxiv.2601.13742,
  title  = {Hearing Between the Lines: Unlocking the Reasoning Power of LLMs for Speech Evaluation},
  author = {Arjun Chandra and Kevin Miller and Venkatesh Ravichandran and Constantinos Papayiannis and Venkatesh Saligrama},
  journal= {arXiv preprint arXiv:2601.13742},
  year   = {2026}
}

Comments

EACL 2026 Findings