中文
相关论文

相关论文: Real-World En Call Center Transcripts Dataset with…

200 篇论文

English is the most widely spoken language in the world, used daily by millions of people as a first or second language in many different contexts. As a result, there are many varieties of English. Although the great many advances in…

计算与语言 · 计算机科学 2023-04-03 Ramon Sanabria , Nikolay Bogoychev , Nina Markl , Andrea Carmantini , Ondrej Klejch , Peter Bell

Model Context Protocol (MCP) enables agents to interact with external tools, yet empirical research on MCP is hindered by the lack of large-scale, accessible datasets. We present MCPZoo, the largest and most comprehensive dataset of MCP…

密码学与安全 · 计算机科学 2025-12-29 Mengying Wu , Pei Chen , Geng Hong , Baichao An , Jinsong Chen , Binwang Wan , Xudong Pan , Jiarun Dai , Min Yang

Recent works demonstrate that voice assistants do not perform equally well for everyone, but research on demographic robustness of speech technologies is still scarce. This is mainly due to the rarity of large datasets with controlled…

Prevalent ungrammatical expressions and disfluencies in spontaneous speech from second language (L2) learners pose unique challenges to Automatic Speech Recognition (ASR) systems. However, few datasets are tailored to L2 learner speech. We…

计算与语言 · 计算机科学 2024-10-07 Haechan Kim , Junho Myung , Seoyoung Kim , Sungpah Lee , Dongyeop Kang , Juho Kim

De-identification of data used for automatic speech recognition modeling is a critical component in protecting privacy, especially in the medical domain. However, simply removing all personally identifiable information (PII) from end-to-end…

音频与语音处理 · 电气工程与系统科学 2022-07-13 Martin Flechl , Shou-Chun Yin , Junho Park , Peter Skala

We present the corpus called HealthCall. This was recorded in real-life conditions in the call center of Malakoff Humanis. It includes two separate audio channels, the first one for the customer and the second one for the agent. Each…

计算与语言 · 计算机科学 2022-08-23 Nikola Lackovic , Claude Montacié , Gauthier Lalande , Marie-José Caraty

The majority of current Text-to-Speech (TTS) datasets, which are collections of individual utterances, contain few conversational aspects. In this paper, we introduce DailyTalk, a high-quality conversational speech dataset designed for…

音频与语音处理 · 电气工程与系统科学 2023-03-14 Keon Lee , Kyumin Park , Daeyoung Kim

Research involving privacy-sensitive data has always been constrained by data scarcity, standing in sharp contrast to other areas that have benefited from data scaling. This challenge is becoming increasingly urgent as modern AI…

We introduce a new collection of spoken English audio suitable for training speech recognition systems under limited or no supervision. It is derived from open-source audio books from the LibriVox project. It contains over 60K hours of…

We present Curated Industrial Developer Repository (CIDR), a large-scale dataset of real-world software repositories collected through direct collaboration with 12 industrial partner organizations. The dataset comprises 2,440 repositories…

软件工程 · 计算机科学 2026-05-13 Vladislav Savenkov

This paper introduces a new large consent-driven dataset aimed at assisting in the evaluation of algorithmic bias and robustness of computer vision and audio speech models in regards to 11 attributes that are self-provided or labeled by…

计算机视觉与模式识别 · 计算机科学 2023-03-10 Bilal Porgali , Vítor Albiero , Jordan Ryda , Cristian Canton Ferrer , Caner Hazirbas

The availability of large, high-quality emotional speech databases is essential for advancing speech emotion recognition (SER) in real-world scenarios. However, many existing databases face limitations in size, emotional balance, and…

We present the first sizeable corpus of Thai speech emotion recognition, THAI-SER, containing 41 hours and 36 minutes (27,854 utterances) from 100 recordings made in different recording environments: Zoom and two studio setups. The…

Africa has a very low doctor-to-patient ratio. At very busy clinics, doctors could see 30+ patients per day -- a heavy patient burden compared with developed countries -- but productivity tools such as clinical automatic speech recognition…

Large-scale multimodal models achieve strong results on tasks like Visual Question Answering (VQA), but they are often limited when queries require cultural and visual information, everyday knowledge, particularly in low-resource and…

The objective of this paper is speaker recognition under noisy and unconstrained conditions. We make two key contributions. First, we introduce a very large-scale audio-visual speaker recognition dataset collected from open-source media.…

声音 · 计算机科学 2020-11-05 Joon Son Chung , Arsha Nagrani , Andrew Zisserman

Recent advances in audio-language models have demonstrated remarkable success on short, segment-level speech tasks. However, real-world applications such as meeting transcription, spoken document understanding, and conversational analysis…

This paper introduces BIRD, the Big Impulse Response Dataset. This open dataset consists of 100,000 multichannel room impulse responses (RIRs) generated from simulations using the Image Method, making it the largest multichannel open…

We describe a system that generates speaker-annotated transcripts of meetings by using a virtual microphone array, a set of spatially distributed asynchronous recording devices such as laptops and mobile phones. The system is composed of…

音频与语音处理 · 电气工程与系统科学 2019-07-09 Takuya Yoshioka , Zhuo Chen , Dimitrios Dimitriadis , William Hinthorn , Xuedong Huang , Andreas Stolcke , Michael Zeng

Everyday speech conveys far more than words, it reflects who we are, how we feel, and the circumstances surrounding our interactions. Yet, most existing speech datasets are acted, limited in scale, and fail to capture the expressive…

音频与语音处理 · 电气工程与系统科学 2025-11-04 Zongyang Du , Shreeram Suresh Chandra , Ismail Rasim Ulgen , Aurosweta Mahapatra , Ali N. Salman , Carlos Busso , Berrak Sisman