English

Best-of-$N$ TTS Evaluation is Confounded by ASR Family Alignment

Computation and Language 2026-07-09 v1 Artificial Intelligence Machine Learning Sound

Abstract

Best-of-NN (BoN) inference improves content consistency in zero-shot text-to-speech by selecting from NN candidates with an automatic speech recognition (ASR) verifier. We identify an underexplored evaluation confound: a verifier's apparent quality depends strongly on which ASR family judges it. On LibriSpeech-PC test-clean~\citep{librispeechpc} with F5-TTS~\citep{f5tts}, verifier rankings reverse across Whisper, wav2vec~2.0, and HuBERT evaluators, and same-family verifier-evaluator pairs recover 2-3×\times more oracle headroom than cross-family pairs despite near-identical representations (linear CKA 0.9780.978) -- a pattern consistent with identity- or lineage-level coupling rather than representational overlap. We propose two \textbf{cross-family rank ensembles} (rank-averaging and conjunctive max-rank) that attain the lowest mean WER across three independent evaluators -- 1.61%1.61\% at N=10N{=}10 (12%-12\% relative to F5-TTS) -- with no measurable degradation under automatic SIM-o/UTMOS metrics; the best single verifier drives WER from 2.06%2.06\% to 1.72%1.72\% (16.5%-16.5\%) under the official F5-TTS evaluator. We recommend cross-evaluator triangulation as default reporting practice.

Keywords

Cite

@article{arxiv.2607.08256,
  title  = {Best-of-$N$ TTS Evaluation is Confounded by ASR Family Alignment},
  author = {Taehyung Yu and Seongjae Kang},
  journal= {arXiv preprint arXiv:2607.08256},
  year   = {2026}
}

Comments

Accepted at ICML 2026 Workshop on Machine Learning for Audio