GigaSpeechBench: A Real-World Multilingual Speech-to-Text Benchmark
Abstract
While modern ASR systems achieve low error rates on high-resource benchmarks, such performance often overestimates real-world robustness. Existing evaluations address challenges in isolation, lacking a unified benchmark for domain terminology, age variation, dialects, accents, and low-resource languages, particularly across the Middle East and Southeast Asia, representing over one billion under-evaluated speakers. To address this gap, we introduce GigaSpeechBench, a comprehensive multilingual and multidimensional in-the-wild ASR & AST benchmark comprising 680 hours of human-annotated speech. It features five modules: (1) 12 low-resource Middle Eastern and Southeast Asian languages, plus challenging Japanese and Korean; (2) 6 Chinese dialects; (3) 6 English accents; (4) dense terminology across 12 vertical domains for Chinese and English; and (5) older adult and child speech. We further provide human-annotated Chinese and English translations for 11 languages to support AST evaluation. Extensive evaluations of leading foundation models and commercial APIs reveal significant performance degradation in these challenging settings, exposing critical evaluation blind spots.
Cite
@article{arxiv.2606.28884,
title = {GigaSpeechBench: A Real-World Multilingual Speech-to-Text Benchmark},
author = {Yujie Tu and Yifan Yang and Tianrui Wang and Yanqiao Zhu and Guodong Lin and Mingchen Shao and Haoran Wang and Junzhe Liu and Yuxiang Fu and Yizhou Peng and Changsong Liu and Peng Wang and Zhikang Niu and Yunchong Xiao and Haolong Zheng and Xiuwen Zheng and Xulin Fan and Wei-Qiang Zhang and Lei Xie and Longbiao Wang and Eng-Siong Chng and Jiajun Zhang and Kele Xu and Jianwei Yu and Binbin Zhang and Jiayu Du and Wupeng Wang and Zhigao Chen and Yunlong Wu and Guoguo Chen and Xipeng Qiu and Mark Hasegawa-Johnson and Kai Yu and Zhifu Gao and Xiangang Li and Xie Chen},
journal= {arXiv preprint arXiv:2606.28884},
year = {2026}
}