NVBench: A Benchmark for Speech Synthesis with Non-Verbal Vocalizations
Abstract
Non-verbal vocalizations (NVVs) like laugh, sigh, and sob are essential for human-like speech, yet standardized evaluation remains limited in jointly assessing whether systems can generate the intended NVVs, place them correctly, and keep them salient without harming speech. We present Non-verbal Vocalization Benchmark (NVBench), a bilingual (English/Chinese) benchmark that evaluates speech synthesis with NVVs. NVBench pairs a unified 45-type taxonomy with a curated bilingual dataset and introduces a multi-axis protocol that separates general speech naturalness and quality from NVV-specific controllability, placement, and salience. We benchmark 15 TTS systems using objective metrics, listening tests, and an LLM-based multi-rater evaluation. Results reveal that NVVs controllability often decouples from quality, while low-SNR oral cues and long-duration affective NVVs remain persistent bottlenecks. NVBench enables fair cross-system comparison across diverse control interfaces under a unified, standardized framework.
Cite
@article{arxiv.2604.16211,
title = {NVBench: A Benchmark for Speech Synthesis with Non-Verbal Vocalizations},
author = {Liumeng Xue and Weizhen Bian and Jiahao Pan and Wenxuan Wang and Yilin Ren and Boyi Kang and Jingbin Hu and Ziyang Ma and Shuai Wang and Xinyuan Qian and Hung-yi Lee and Yike Guo},
journal= {arXiv preprint arXiv:2604.16211},
year = {2026}
}