Seed-ASR: Understanding Diverse Speech and Contexts with LLM-based Speech Recognition
Abstract
Modern automatic speech recognition (ASR) model is required to accurately transcribe diverse speech signals (from different domains, languages, accents, etc) given the specific contextual information in various application scenarios. Classic end-to-end models fused with extra language models perform well, but mainly in data matching scenarios and are gradually approaching a bottleneck. In this work, we introduce Seed-ASR, a large language model (LLM) based speech recognition model. Seed-ASR is developed based on the framework of audio conditioned LLM (AcLLM), leveraging the capabilities of LLMs by inputting continuous speech representations together with contextual information into the LLM. Through stage-wise large-scale training and the elicitation of context-aware capabilities in LLM, Seed-ASR demonstrates significant improvement over end-to-end models on comprehensive evaluation sets, including multiple domains, accents/dialects and languages. Additionally, Seed-ASR can be further deployed to support specific needs in various scenarios without requiring extra language models. Compared to recently released large ASR models, Seed-ASR achieves 10%-40% reduction in word (or character, for Chinese) error rates on Chinese and English public test sets, further demonstrating its powerful performance.
Keywords
Cite
@article{arxiv.2407.04675,
title = {Seed-ASR: Understanding Diverse Speech and Contexts with LLM-based Speech Recognition},
author = {Ye Bai and Jingping Chen and Jitong Chen and Wei Chen and Zhuo Chen and Chuang Ding and Linhao Dong and Qianqian Dong and Yujiao Du and Kepan Gao and Lu Gao and Yi Guo and Minglun Han and Ting Han and Wenchao Hu and Xinying Hu and Yuxiang Hu and Deyu Hua and Lu Huang and Mingkun Huang and Youjia Huang and Jishuo Jin and Fanliu Kong and Zongwei Lan and Tianyu Li and Xiaoyang Li and Zeyang Li and Zehua Lin and Rui Liu and Shouda Liu and Lu Lu and Yizhou Lu and Jingting Ma and Shengtao Ma and Yulin Pei and Chen Shen and Tian Tan and Xiaogang Tian and Ming Tu and Bo Wang and Hao Wang and Yuping Wang and Yuxuan Wang and Hanzhang Xia and Rui Xia and Shuangyi Xie and Hongmin Xu and Meng Yang and Bihong Zhang and Jun Zhang and Wanyi Zhang and Yang Zhang and Yawei Zhang and Yijie Zheng and Ming Zou},
journal= {arXiv preprint arXiv:2407.04675},
year = {2024}
}