English
Related papers

Related papers: AISHELL-4: An Open Source Dataset for Speech Enhan…

200 papers

This paper delineates AISHELL-5, the first open-source in-car multi-channel multi-speaker Mandarin automatic speech recognition (ASR) dataset. AISHLL-5 includes two parts: (1) over 100 hours of multi-channel speech data recorded in an…

Sound · Computer Science 2025-05-30 Yuhang Dai , He Wang , Xingchen Li , Zihan Zhang , Shuiyuan Wang , Lei Xie , Xin Xu , Hongxiao Guo , Shaoji Zhang , Hui Bu , Wei Chen

AISHELL-1 is by far the largest open-source speech corpus available for Mandarin speech recognition research. It was released with a baseline system containing solid training and testing pipelines for Mandarin ASR. In AISHELL-2, 1000 hours…

Computation and Language · Computer Science 2018-09-14 Jiayu Du , Xingyu Na , Xuechen Liu , Hui Bu

In this paper, we present AISHELL-3, a large-scale and high-fidelity multi-speaker Mandarin speech corpus which could be used to train multi-speaker Text-to-Speech (TTS) systems. The corpus contains roughly 85 hours of emotion-neutral…

Sound · Computer Science 2021-04-23 Yao Shi , Hui Bu , Xin Xu , Shaoji Zhang , Ming Li

Whisper speech recognition is crucial not only for ensuring privacy in sensitive communications but also for providing a critical communication bridge for patients under vocal restraint and enabling discrete interaction in noise-sensitive…

Audio and Speech Processing · Electrical Eng. & Systems 2025-09-30 Cancan Li , Fei Su , Juan Liu , Hui Bu , Yulong Wan , Hongbin Suo , Ming Li

An open-source Mandarin speech corpus called AISHELL-1 is released. It is by far the largest corpus which is suitable for conducting the speech recognition research and building speech recognition systems for Mandarin. The recording…

Computation and Language · Computer Science 2017-09-19 Hui Bu , Jiayu Du , Xingyu Na , Bengu Wu , Hao Zheng

Recent development of speech processing, such as speech recognition, speaker diarization, etc., has inspired numerous applications of speech technologies. The meeting scenario is one of the most valuable and, at the same time, most…

Sound · Computer Science 2022-02-28 Fan Yu , Shiliang Zhang , Yihui Fu , Lei Xie , Siqi Zheng , Zhihao Du , Weilong Huang , Pengcheng Guo , Zhijie Yan , Bin Ma , Xin Xu , Hui Bu

This paper introduces a new open-sourced Mandarin speech corpus, called DiDiSpeech. It consists of about 800 hours of speech data at 48kHz sampling rate from 6000 speakers and the corresponding texts. All speech data in the corpus is…

Audio and Speech Processing · Electrical Eng. & Systems 2021-02-09 Tingwei Guo , Cheng Wen , Dongwei Jiang , Ne Luo , Ruixiong Zhang , Shuaijiang Zhao , Wubo Li , Cheng Gong , Wei Zou , Kun Han , Xiangang Li

Code-switching (CS), the alternation between two or more languages within a single conversation, presents significant challenges for automatic speech recognition (ASR) systems. Existing Mandarin-English code-switching datasets often suffer…

Computation and Language · Computer Science 2025-03-13 Jiaming Zhou , Yujie Guo , Shiwan Zhao , Haoqin Sun , Hui Wang , Jiabei He , Aobo Kong , Shiyao Wang , Xi Yang , Yequan Wang , Yonghua Lin , Yong Qin

The ICASSP 2022 Multi-channel Multi-party Meeting Transcription Grand Challenge (M2MeT) focuses on one of the most valuable and the most challenging scenarios of speech technologies. The M2MeT challenge has particularly set up two tracks,…

Automatic speech recognition (ASR) systems have advanced significantly with models like Whisper, Conformer, and self-supervised frameworks such as Wav2vec 2.0 and HuBERT. However, developing robust ASR models for young children's speech…

Speaker diarization(SD) is a classic task in speech processing and is crucial in multi-party scenarios such as meetings and conversations. Current mainstream speaker diarization approaches consider acoustic information only, which result in…

Computation and Language · Computer Science 2023-05-23 Luyao Cheng , Siqi Zheng , Zhang Qinglin , Hui Wang , Yafeng Chen , Qian Chen

The evolving speech processing landscape is increasingly focused on complex scenarios like meetings or cocktail parties with multiple simultaneous speakers and far-field conditions. Existing methodologies for addressing these challenges…

We introduce the first Natural Office Talkers in Settings of Far-field Audio Recordings (``NOTSOFAR-1'') Challenge alongside datasets and baseline system. The challenge focuses on distant speaker diarization and automatic speech recognition…

This paper introduces a high-quality rich annotated Mandarin conversational (RAMC) speech dataset called MagicData-RAMC. The MagicData-RAMC corpus contains 180 hours of conversational speech data recorded from native speakers of Mandarin…

Computation and Language · Computer Science 2022-04-01 Zehui Yang , Yifan Chen , Lei Luo , Runyan Yang , Lingxuan Ye , Gaofeng Cheng , Ji Xu , Yaohui Jin , Qingqing Zhang , Pengyuan Zhang , Lei Xie , Yonghong Yan

This paper describes our speaker diarization system submitted to the Multi-channel Multi-party Meeting Transcription (M2MeT) challenge, where Mandarin meeting data were recorded in multi-channel format for diarization and automatic speech…

Audio and Speech Processing · Electrical Eng. & Systems 2022-02-07 Naijun Zheng , Na Li , Xixin Wu , Lingwei Meng , Jiawen Kang , Haibin Wu , Chao Weng , Dan Su , Helen Meng

Speaker diarization is the task of partitioning audio into segments according to speaker identity, answering the question of "who spoke when" in multi-speaker conversation recordings. While diarization is an essential task for many…

Sound · Computer Science 2025-10-01 Luca A. Lanzendörfer , Florian Grötschla , Cesare Blaser , Roger Wattenhofer

In this paper, we present WenetSpeech, a multi-domain Mandarin corpus consisting of 10000+ hours high-quality labeled speech, 2400+ hours weakly labeled speech, and about 10000 hours unlabeled speech, with 22400+ hours in total. We collect…

OleSpeech-IV dataset is a large-scale multispeaker and multilingual conversational speech dataset with diverse topics. The audio content comes from publicly-available English podcasts, talk shows, teleconferences, and other conversations.…

Computation and Language · Computer Science 2025-09-08 Wei Chu , Yuanzhe Dong , Ke Tan , Dong Han , Xavier Menendez-Pidal , Ruchao Fan , Chenfeng Miao , Chanwoo Kim , Bhiksha Raj , Rita Singh

Recent advances in audio-language models have demonstrated remarkable success on short, segment-level speech tasks. However, real-world applications such as meeting transcription, spoken document understanding, and conversational analysis…

‹ Prev 1 2 3 10 Next ›