English
Related papers

Related papers: Earnings-21: A Practical Benchmark for ASR in the …

200 papers

We propose a semi-supervised learning method for building end-to-end rich transcription-style automatic speech recognition (RT-ASR) systems from small-scale rich transcription-style and large-scale common transcription-style datasets. In…

Computation and Language · Computer Science 2021-07-13 Tomohiro Tanaka , Ryo Masumura , Mana Ihori , Akihiko Takashima , Shota Orihashi , Naoki Makishima

This paper presents a comprehensive evaluation of Urdu Automatic Speech Recognition (ASR) models. We analyze the performance of three ASR model families: Whisper, MMS, and Seamless-M4T using Word Error Rate (WER), along with a detailed…

Computation and Language · Computer Science 2025-06-09 Samee Arif , Sualeha Farid , Aamina Jamal Khan , Mustafa Abbas , Agha Ali Raza , Awais Athar

The challenge of fairness arises when Automatic Speech Recognition (ASR) systems do not perform equally well for all sub-groups of the population. In the past few years there have been many improvements in overall speech recognition…

Sound · Computer Science 2023-06-12 Irina-Elena Veliche , Pascale Fung

The practical deployment of Audio-Visual Speech Recognition (AVSR) systems is fundamentally challenged by significant performance degradation in real-world environments, characterized by unpredictable acoustic noise and visual interference.…

Audio and Speech Processing · Electrical Eng. & Systems 2025-12-17 Sungnyun Kim

This paper describes an English audio and textual dataset of debating speeches, a unique resource for the growing research field of computational argumentation and debating technologies. We detail the process of speech recording by…

Computation and Language · Computer Science 2018-03-28 Shachar Mirkin , Michal Jacovi , Tamar Lavee , Hong-Kwang Kuo , Samuel Thomas , Leslie Sager , Lili Kotlerman , Elad Venezian , Noam Slonim

Streaming end-to-end automatic speech recognition (ASR) models are widely used on smart speakers and on-device applications. Since these models are expected to transcribe speech with minimal latency, they are constrained to be causal with…

Named entity recognition (NER) is among SLU tasks that usually extract semantic information from textual documents. Until now, NER from speech is made through a pipeline process that consists in processing first an automatic speech…

Computation and Language · Computer Science 2018-05-31 Sahar Ghannay , Antoine Caubrière , Yannick Estève , Antoine Laurent , Emmanuel Morin

The "Switchboard benchmark" is a very well-known test set in automatic speech recognition (ASR) research, establishing record-setting performance for systems that claim human-level transcription accuracy. This work highlights lesser-known…

Audio and Speech Processing · Electrical Eng. & Systems 2022-06-28 Arlo Faria , Adam Janin , Korbinian Riedhammer , Sidhi Adkoli

Speech technologies are transforming interactions across various sectors, from healthcare to call centers and robots, yet their performance on African-accented conversations remains underexplored. We introduce Afrispeech-Dialog, a benchmark…

Speaker adaptation, which involves cloning voices from unseen speakers in the Text-to-Speech task, has garnered significant interest due to its numerous applications in multi-media fields. Despite recent advancements, existing methods often…

Audio and Speech Processing · Electrical Eng. & Systems 2024-07-09 Ruibo Fu , Xin Qi , Zhengqi Wen , Jianhua Tao , Tao Wang , Chunyu Qiang , Zhiyong Wang , Yi Lu , Xiaopeng Wang , Shuchen Shi , Yukun Liu , Xuefei Liu , Shuai Zhang

Recent research using pre-trained transformer models suggests that just 10 minutes of transcribed speech may be enough to fine-tune such a model for automatic speech recognition (ASR) -- at least if we can also leverage vast amounts of text…

Computation and Language · Computer Science 2023-02-13 Nay San , Martijn Bartelds , Blaine Billings , Ella de Falco , Hendi Feriza , Johan Safri , Wawan Sahrozi , Ben Foley , Bradley McDonnell , Dan Jurafsky

The development of Automatic Speech Recognition (ASR) systems for low-resource African languages remains challenging due to limited transcribed speech data. While recent advances in large multilingual models like OpenAI's Whisper offer…

Computation and Language · Computer Science 2025-10-09 Benjamin Akera , Evelyn Nafula , Patrick Walukagga , Gilbert Yiga , John Quinn , Ernest Mwebaze

Sequence-to-sequence models have been widely used in end-to-end speech processing, for example, automatic speech recognition (ASR), speech translation (ST), and text-to-speech (TTS). This paper focuses on an emergent sequence-to-sequence…

This paper describes the systems developed by the HCCL team for the NIST 2021 speaker recognition evaluation (NIST SRE21).We first explore various state-of-the-art speaker embedding extractors combined with a novel circle loss to obtain…

Sound · Computer Science 2022-07-12 Zhuo Li , Runqiu Xiao , Hangting Chen , Zhenduo Zhao , Zihan Zhang , Wenchao Wang

Capitalization and punctuation are important cues for comprehending written texts and conversational transcripts. Yet, many ASR systems do not produce punctuated and case-formatted speech transcripts. We propose to use a multi-task system…

Computation and Language · Computer Science 2021-09-14 Raghavendra Pappagari , Piotr Żelasko , Agnieszka Mikołajczyk , Piotr Pęzik , Najim Dehak

Neural network models for audio tasks, such as automatic speech recognition (ASR) and acoustic scene classification (ASC), are susceptible to noise contamination for real-life applications. To improve audio quality, an enhancement module,…

Bootstrapping speech recognition on limited data resources has been an area of active research for long. The recent transition to all-neural models and end-to-end (E2E) training brought along particular challenges as these models are known…

Audio and Speech Processing · Electrical Eng. & Systems 2021-06-21 Manuel Giollo , Deniz Gunceler , Yulan Liu , Daniel Willett

Neural speaker diarization is widely used for overlap-aware speaker diarization, but it requires large multi-speaker datasets for training. To meet this data requirement, large datasets are often constructed by combining multiple corpora,…

Audio and Speech Processing · Electrical Eng. & Systems 2025-08-26 Shota Horiguchi , Naohiro Tawara , Takanori Ashihara , Atsushi Ando , Marc Delcroix

Despite the tremendous success of automatic speech recognition (ASR) with the introduction of deep learning, its performance is still unsatisfactory in many real-world multi-talker scenarios. Speaker separation excels in separating…

Sound · Computer Science 2025-03-25 Yufeng Yang , Hassan Taherian , Vahid Ahmadi Kalkhorani , DeLiang Wang

Social robots deployed in public spaces present a challenging task for ASR because of a variety of factors, including noise SNR of 20 to 5 dB. Existing ASR models perform well for higher SNRs in this range, but degrade considerably with…

Audio and Speech Processing · Electrical Eng. & Systems 2020-01-15 Charles Jankowski , Vishwas Mruthyunjaya , Ruixi Lin