English
Related papers

Related papers: Real-World En Call Center Transcripts Dataset with…

200 papers

The adaptation of Large-Scale Language Models (LLMs) to specific domains depends on high-quality fine-tuning datasets, particularly in instructional format (e.g., Question-Answer - Q&A). However, generating these datasets, particularly from…

Machine Learning · Computer Science 2026-01-22 Alex Echeverria , Sávio Salvarino Teles de Oliveira , Fernando Marques Federson

This paper introduces a new open-sourced Mandarin speech corpus, called DiDiSpeech. It consists of about 800 hours of speech data at 48kHz sampling rate from 6000 speakers and the corresponding texts. All speech data in the corpus is…

Audio and Speech Processing · Electrical Eng. & Systems 2021-02-09 Tingwei Guo , Cheng Wen , Dongwei Jiang , Ne Luo , Ruixiong Zhang , Shuaijiang Zhao , Wubo Li , Cheng Gong , Wei Zou , Kun Han , Xiangang Li

Speech emotion recognition is a vital contributor to the next generation of human-computer interaction (HCI). However, current existing small-scale databases have limited the development of related research. In this paper, we present LSSED,…

Sound · Computer Science 2021-02-04 Weiquan Fan , Xiangmin Xu , Xiaofen Xing , Weidong Chen , Dongyan Huang

We release the USC Long Single-Speaker (LSS) dataset containing real-time MRI video of the vocal tract dynamics and simultaneous audio obtained during speech production. This unique dataset contains roughly one hour of video and audio data…

We present an open-source speech corpus for the Kazakh language. The Kazakh speech corpus (KSC) contains around 332 hours of transcribed audio comprising over 153,000 utterances spoken by participants from different regions and age groups,…

Audio and Speech Processing · Electrical Eng. & Systems 2021-07-22 Yerbolat Khassanov , Saida Mussakhojayeva , Almas Mirzakhmetov , Alen Adiyev , Mukhamet Nurpeiissov , Huseyin Atakan Varol

Building reliable conversational AI assistants for customer-facing industries remains challenging due to noisy conversational data, fragmented knowledge, and the requirement for accurate human hand-off - particularly in domains that depend…

Computation and Language · Computer Science 2026-02-19 Krittin Pachtrachai , Petmongkon Pornpichitsuwan , Wachiravit Modecrua , Touchapon Kraisingkorn

This paper presents BSTC (Baidu Speech Translation Corpus), a large-scale Chinese-English speech translation dataset. This dataset is constructed based on a collection of licensed videos of talks or lectures, including about 68 hours of…

Computation and Language · Computer Science 2021-04-28 Ruiqing Zhang , Xiyang Wang , Chuanqiang Zhang , Zhongjun He , Hua Wu , Zhi Li , Haifeng Wang , Ying Chen , Qinfei Li

This paper delineates AISHELL-5, the first open-source in-car multi-channel multi-speaker Mandarin automatic speech recognition (ASR) dataset. AISHLL-5 includes two parts: (1) over 100 hours of multi-channel speech data recorded in an…

Sound · Computer Science 2025-05-30 Yuhang Dai , He Wang , Xingchen Li , Zihan Zhang , Shuiyuan Wang , Lei Xie , Xin Xu , Hongxiao Guo , Shaoji Zhang , Hui Bu , Wei Chen

Engagement in virtual learning is essential for participant satisfaction, performance, and adherence, particularly in online education and virtual rehabilitation, where interactive communication plays a key role. Yet, accurately measuring…

Computer Vision and Pattern Recognition · Computer Science 2025-07-25 Ali Abedi , Sadaf Safa , Tracey J. F. Colella , Shehroz S. Khan

Recent advancements in speech generation models have been significantly driven by the use of large-scale training data. However, producing highly spontaneous, human-like speech remains a challenge due to the scarcity of large, diverse, and…

Audio and Speech Processing · Electrical Eng. & Systems 2024-09-10 Haorui He , Zengqiang Shang , Chaoren Wang , Xuyuan Li , Yicheng Gu , Hua Hua , Liwei Liu , Chen Yang , Jiaqi Li , Peiyang Shi , Yuancheng Wang , Kai Chen , Pengyuan Zhang , Zhizheng Wu

Talk2AI is a large-scale longitudinal dataset of 3,080 conversations (totaling 30,800 turns) between human participants and Large Language Models (LLMs), designed to support research on persuasion, opinion change, and human-AI interaction.…

Human-Computer Interaction · Computer Science 2026-04-07 Alexis Carrillo , Enrique Taietta , Ali Aghazadeh Ardebili , Giuseppe Alessandro Veltri , Massimo Stella

Recently, researchers set an ambitious goal of conducting speaker recognition in unconstrained conditions where the variations on ambient, channel and emotion could be arbitrary. However, most publicly available datasets are collected under…

Audio and Speech Processing · Electrical Eng. & Systems 2019-11-06 Yue Fan , Jiawen Kang , Lantian Li , Kaicheng Li , Haolin Chen , Sitong Cheng , Pengyuan Zhang , Ziya Zhou , Yunqi Cai , Dong Wang

This document provides a brief description of the National Institute of Standards and Technology (NIST) speaker recognition evaluation (SRE) conversational telephone speech (CTS) Superset. The CTS Superset has been created in an attempt to…

Sound · Computer Science 2021-08-17 Seyed Omid Sadjadi

The Common Voice corpus is a massively-multilingual collection of transcribed speech intended for speech technology research and development. Common Voice is designed for Automatic Speech Recognition purposes but can be useful in other…

aTrain is an open-source and offline tool for transcribing audio data in multiple languages with CPU and NVIDIA GPU support. It is specifically designed for researchers using qualitative data generated from various forms of speech…

Sound · Computer Science 2023-10-19 Armin Haberl , Jürgen Fleiß , Dominik Kowald , Stefan Thalmann

This paper introduces a new multi-speaker English dataset for training text-to-speech models. The dataset is based on LibriVox audiobooks and Project Gutenberg texts, both in the public domain. The new dataset contains about 292 hours of…

Audio and Speech Processing · Electrical Eng. & Systems 2021-06-16 Evelina Bakhturina , Vitaly Lavrukhin , Boris Ginsburg , Yang Zhang

Speech translation for Indian languages remains a challenging task due to the scarcity of large-scale, publicly available datasets that capture the linguistic diversity and domain coverage essential for real-world applications. Existing…

Despite advancements in conversational AI, language models encounter challenges to handle diverse conversational tasks, and existing dialogue dataset collections often lack diversity and comprehensiveness. To tackle these issues, we…

Computation and Language · Computer Science 2024-02-06 Jianguo Zhang , Kun Qian , Zhiwei Liu , Shelby Heinecke , Rui Meng , Ye Liu , Zhou Yu , Huan Wang , Silvio Savarese , Caiming Xiong

This paper presents BUT ReverbDB - a dataset of real room impulse responses (RIR), background noises and re-transmitted speech data. The retransmitted data includes LibriSpeech test-clean, 2000 HUB5 English evaluation and part of 2010 NIST…

Audio and Speech Processing · Electrical Eng. & Systems 2019-09-04 Igor Szoke , Miroslav Skacel , Ladislav Mosner , Jakub Paliesek , Jan "Honza" Cernocky

Studying how people interact with large language models (LLMs) in real-world scenarios is increasingly important due to their widespread use in various applications. In this paper, we introduce LMSYS-Chat-1M, a large-scale dataset…