中文
相关论文

相关论文: WenetSpeech: A 10000+ Hours Multi-domain Mandarin …

200 篇论文

Dubbed series are gaining a lot of popularity in recent years with strong support from major media service providers. Such popularity is fueled by studies that showed that dubbed versions of TV shows are more popular than their subtitled…

计算与语言 · 计算机科学 2022-03-08 Massa Baali , Wassim El-Hajj , Ahmed Ali

In this paper, we introduce a high-quality and large-scale benchmark dataset for English-Vietnamese speech translation with 508 audio hours, consisting of 331K triplets of (sentence-lengthed audio, English source transcript sentence,…

计算与语言 · 计算机科学 2022-08-09 Linh The Nguyen , Nguyen Luong Tran , Long Doan , Manh Luong , Dat Quoc Nguyen

Streaming end-to-end automatic speech recognition (ASR) models are widely used on smart speakers and on-device applications. Since these models are expected to transcribe speech with minimal latency, they are constrained to be causal with…

Closed-Set speaker identification aims to assign a speech utterance to one of a predefined set of enrolled speakers and requires robust modeling of speaker-specific characteristics across multiple temporal scales. While recent deep learning…

声音 · 计算机科学 2026-05-11 Yassin Terraf , Youssef Iraqi

Recent progress in speech processing has highlighted that high-quality performance across languages requires substantial training data for each individual language. While existing multilingual datasets cover many languages, they often…

计算与语言 · 计算机科学 2025-10-28 Samuel Pfisterer , Florian Grötschla , Luca A. Lanzendörfer , Florian Yan , Roger Wattenhofer

Self-supervised learning (SSL) is at the origin of unprecedented improvements in many different domains including computer vision and natural language processing. Speech processing drastically benefitted from SSL as most of the current…

We present the LEMAS-Dataset, which, to our knowledge, is currently the largest open-source multilingual speech corpus with word-level timestamps. Covering over 150,000 hours across 10 major languages, LEMAS-Dataset is constructed via a…

声音 · 计算机科学 2026-01-09 Zhiyuan Zhao , Lijian Lin , Ye Zhu , Kai Xie , Yunfei Liu , Yu Li

Benchmarking plays a pivotal role in assessing and enhancing the performance of compact deep learning models designed for execution on resource-constrained devices, such as microcontrollers. Our study introduces a novel, entirely…

声音 · 计算机科学 2024-03-18 René Groh , Nina Goes , Andreas M. Kist

This paper presents AMNet, an Acoustic Model Network designed to improve the performance of Mandarin speech synthesis by incorporating phrase structure annotation and local convolution modules. AMNet builds upon the FastSpeech 2…

声音 · 计算机科学 2025-04-15 Yubing Cao , Yinfeng Yu , Yongming Li , Liejun Wang

Even for better-studied sign languages like American Sign Language (ASL), data is the bottleneck for machine learning research. The situation is worse yet for the many other sign languages used by Deaf/Hard of Hearing communities around the…

计算与语言 · 计算机科学 2024-07-17 Garrett Tanzer , Biao Zhang

The training of deep learning-based multichannel speech enhancement and source localization systems relies heavily on the simulation of room impulse response and multichannel diffuse noise, due to the lack of large-scale real-recorded…

声音 · 计算机科学 2024-10-02 Bing Yang , Changsheng Quan , Yabo Wang , Pengyu Wang , Yujie Yang , Ying Fang , Nian Shao , Hui Bu , Xin Xu , Xiaofei Li

This paper introduces RyanSpeech, a new speech corpus for research on automated text-to-speech (TTS) systems. Publicly available TTS corpora are often noisy, recorded with multiple speakers, or lack quality male speech data. In order to…

计算与语言 · 计算机科学 2021-06-17 Rohola Zandie , Mohammad H. Mahoor , Julia Madsen , Eshrat S. Emamian

We present the Multilingual TEDx corpus, built to support speech recognition (ASR) and speech translation (ST) research across many non-English source languages. The corpus is a collection of audio recordings from TEDx talks in 8 source…

We present INDICVOICES, a dataset of natural and spontaneous speech containing a total of 7348 hours of read (9%), extempore (74%) and conversational (17%) audio from 16237 speakers covering 145 Indian districts and 22 languages. Of these…

Automatic speech recognition (ASR) in clinical dialogue demands robustness to full-duplex interaction, speaker overlap, and low-latency constraints, yet open benchmarks remain scarce. We present MMedFD, the first real-world Chinese…

音频与语音处理 · 电气工程与系统科学 2025-09-29 Hongzhao Chen , XiaoYang Wang , Jing Lan , Hexiao Ding , Yufeng Jiang , MingHui Yang , DanHui Xu , Jun Luo , Nga-Chun Ng , Gerald W. Y. Cheng , Yunlin Mao , Jung Sun Yoo

Active speaker detection is an important component in video analysis algorithms for applications such as speaker diarization, video re-targeting for meetings, speech enhancement, and human-robot interaction. The absence of a large,…

As increasing development of text-to-speech (TTS) and voice conversion (VC) technologies, the detection of synthetic speech has been suffered dramatically. In order to promote the development of synthetic speech detection model against…

声音 · 计算机科学 2021-10-19 Zhenyu Zhang , Yewei Gu , Xiaowei Yi , Xianfeng Zhao

The performance of speech-processing models is heavily influenced by the speech corpus that is used for training and evaluation. In this study, we propose BAlanced Script PROducer (BASPRO) system, which can automatically construct a…

神经与进化计算 · 计算机科学 2023-01-11 Yu-Wen Chen , Hsin-Min Wang , Yu Tsao

This paper presents a high quality Vietnamese speech corpus that can be used for analyzing Vietnamese speech characteristic as well as building speech synthesis models. The corpus consists of 5400 clean-speech utterances spoken by 12…

计算与语言 · 计算机科学 2019-04-12 Pham Ngoc Phuong , Quoc Truong Do , Luong Chi Mai

We present an open-source speech corpus for the Kazakh language. The Kazakh speech corpus (KSC) contains around 332 hours of transcribed audio comprising over 153,000 utterances spoken by participants from different regions and age groups,…

音频与语音处理 · 电气工程与系统科学 2021-07-22 Yerbolat Khassanov , Saida Mussakhojayeva , Almas Mirzakhmetov , Alen Adiyev , Mukhamet Nurpeiissov , Huseyin Atakan Varol