English
Related papers

Related papers: Real-World En Call Center Transcripts Dataset with…

200 papers

Existing conversational datasets consist either of written proxies for dialog or small-scale transcriptions of natural speech. We introduce 'Interview': a large-scale (105K conversations) media dialog dataset collected from news interview…

Computation and Language · Computer Science 2020-04-08 Bodhisattwa Prasad Majumder , Shuyang Li , Jianmo Ni , Julian McAuley

Evaluating English ASR systems for conversational AI applications remains difficult, as many publicly available corpora are either pre-segmented into short segments, consist of read or prepared speech, or lack explicit dialect annotations…

Computation and Language · Computer Science 2026-05-01 Eugen Beck , Sarah Beranek , Uma Moothiringote , Daniel Mann , Wilfried Michel , Katie Nguyen , Taylor Tragemann

Existing Indic ASR benchmarks often use scripted, clean speech and leaderboard driven evaluation that encourages dataset specific overfitting. In addition, strict single reference WER penalizes natural spelling variation in Indian…

We present INDICVOICES, a dataset of natural and spontaneous speech containing a total of 7348 hours of read (9%), extempore (74%) and conversational (17%) audio from 16237 speakers covering 145 Indian districts and 22 languages. Of these…

In this paper, we present AISHELL-4, a sizable real-recorded Mandarin speech dataset collected by 8-channel circular microphone array for speech processing in conference scenario. The dataset consists of 211 recorded meeting sessions, each…

Sound · Computer Science 2021-08-11 Yihui Fu , Luyao Cheng , Shubo Lv , Yukai Jv , Yuxiang Kong , Zhuo Chen , Yanxin Hu , Lei Xie , Jian Wu , Hui Bu , Xin Xu , Jun Du , Jingdong Chen

This paper introduces a new corpus of Mandarin-English code-switching speech recognition--TALCS corpus, suitable for training and evaluating code-switching speech recognition systems. TALCS corpus is derived from real online one-to-one…

Computation and Language · Computer Science 2022-06-28 Chengfei Li , Shuhao Deng , Yaoping Wang , Guangjing Wang , Yaguang Gong , Changbin Chen , Jinfeng Bai

Medical dialogue systems are promising in assisting in telemedicine to increase access to healthcare services, improve the quality of patient care, and reduce medical costs. To facilitate the research and development of medical dialogue…

Recent advancements in speech generation have been driven by large-scale training datasets. However, current models struggle to capture the spontaneity and variability inherent in real-world human speech, as they are primarily trained on…

Bengali is one of the most spoken languages in the world with over 300 million speakers globally. Despite its popularity, research into the development of Bengali speech recognition systems is hindered due to the lack of diverse open-source…

The Speech Wikimedia Dataset is a publicly available compilation of audio with transcriptions extracted from Wikimedia Commons. It includes 1780 hours (195 GB) of CC-BY-SA licensed transcribed speech from a diverse set of scenarios and…

Artificial Intelligence · Computer Science 2023-08-31 Rafael Mosquera Gómez , Julián Eusse , Juan Ciro , Daniel Galvez , Ryan Hileman , Kurt Bollacker , David Kanter

Automatic speech recognition (ASR) research is driven by the availability of common datasets between industrial researchers and academics, encouraging comparisons and evaluations. LibriSpeech, despite its long success as an ASR benchmark,…

Computation and Language · Computer Science 2025-05-29 Titouan Parcollet , Yuan Tseng , Shucong Zhang , Rogier van Dalen

Automatic speech recognition (ASR) performs well for high-resource languages with abundant paired audio-transcript data, but its accuracy degrades sharply for most languages due to limited publicly available aligned data. To this end, we…

Computation and Language · Computer Science 2026-05-12 Antonis Asonitis , Luca A. Lanzendörfer , Frédéric Berdoz , Roger Wattenhofer

Code-switching (CS), the alternation between two or more languages within a single conversation, presents significant challenges for automatic speech recognition (ASR) systems. Existing Mandarin-English code-switching datasets often suffer…

Computation and Language · Computer Science 2025-03-13 Jiaming Zhou , Yujie Guo , Shiwan Zhao , Haoqin Sun , Hui Wang , Jiabei He , Aobo Kong , Shiyao Wang , Xi Yang , Yequan Wang , Yonghua Lin , Yong Qin

A random sample of nearly 10 hours of speech from PennSound, the world's largest online collection of poetry readings and discussions, was used as a benchmark to evaluate several commercial and open-source speech-to-text systems.…

Computation and Language · Computer Science 2025-04-09 Jonathan Wright , Mark Liberman , Neville Ryant , James Fiumara

The current public datasets for speech recognition (ASR) tend not to focus specifically on the fairness aspect, such as performance across different demographic groups. This paper introduces a novel dataset, Fair-Speech, a publicly released…

Artificial Intelligence · Computer Science 2024-08-26 Irina-Elena Veliche , Zhuangqun Huang , Vineeth Ayyat Kochaniyan , Fuchun Peng , Ozlem Kalinli , Michael L. Seltzer

It is well known that many machine learning systems demonstrate bias towards specific groups of individuals. This problem has been studied extensively in the Facial Recognition area, but much less so in Automatic Speech Recognition (ASR).…

Audio and Speech Processing · Electrical Eng. & Systems 2021-11-22 Chunxi Liu , Michael Picheny , Leda Sarı , Pooja Chitkara , Alex Xiao , Xiaohui Zhang , Mark Chou , Andres Alvarado , Caner Hazirbas , Yatharth Saraf

This paper introduces the Voices Obscured In Complex Environmental Settings (VOICES) corpus, a freely available dataset under Creative Commons BY 4.0. This dataset will promote speech and signal processing research of speech recorded by…

The People's Speech is a free-to-download 30,000-hour and growing supervised conversational English speech recognition dataset licensed for academic and commercial usage under CC-BY-SA (with a CC-BY subset). The data is collected via…

Speech Emotion recognition (SER) in call center conversations has emerged as a valuable tool for assessing the quality of interactions between clients and agents. In contrast to controlled laboratory environments, real-life conversations…

Audio and Speech Processing · Electrical Eng. & Systems 2023-10-05 Yajing Feng , Laurence Devillers

Recent advances in Automatic Speech Recognition (ASR) have made it possible to reliably produce automatic transcripts of clinician-patient conversations. However, access to clinical datasets is heavily restricted due to patient privacy,…

Computation and Language · Computer Science 2022-04-04 Alex Papadopoulos Korfiatis , Francesco Moramarco , Radmila Sarac , Aleksandar Savkov
‹ Prev 1 2 3 10 Next ›