中文
相关论文

相关论文: DiPCo -- Dinner Party Corpus

200 篇论文

We introduce a new dataset of conversational speech representing English from India, Nigeria, and the United States. The Multi-Dialect Dataset of Dialogues (MD3) strikes a new balance between open-ended conversational speech and…

计算与语言 · 计算机科学 2023-05-22 Jacob Eisenstein , Vinodkumar Prabhakaran , Clara Rivera , Dorottya Demszky , Devyani Sharma

The uptake of deep learning in natural language generation (NLG) led to the release of both small and relatively large parallel corpora for training neural models. The existing data-to-text datasets are, however, aimed at task-oriented…

计算与语言 · 计算机科学 2019-10-29 Juraj Juraska , Kevin K. Bowden , Marilyn Walker

We introduce TiCo, a time-controllable spoken dialogue model (SDM) that follows time-constrained instructions (e.g., "Please generate a response lasting about 15 seconds") and generates spoken responses with controllable duration. This…

计算与语言 · 计算机科学 2026-05-14 Kai-Wei Chang , Wei-Chih Chen , En-Pei Hu , Hung-yi Lee , James Glass

In this study, we present a speech corpus of patients with chronic kidney disease (CKD) that will be used for research on pathological voice analysis, automatic illness identification, and severity prediction. This paper introduces the…

计算与语言 · 计算机科学 2022-11-04 Jihyun Mun , Sunhee Kim , Myeong Ju Kim , Jiwon Ryu , Sejoong Kim , Minhwa Chung

The evolving speech processing landscape is increasingly focused on complex scenarios like meetings or cocktail parties with multiple simultaneous speakers and far-field conditions. Existing methodologies for addressing these challenges…

This paper presents the InScript corpus (Narrative Texts Instantiating Script structure). InScript is a corpus of 1,000 stories centered around 10 different scenarios. Verbs and noun phrases are annotated with event and participant types,…

计算与语言 · 计算机科学 2017-03-16 Ashutosh Modi , Tatjana Anikina , Simon Ostermann , Manfred Pinkal

We propose an algorithm to separate simultaneously speaking persons from each other, the "cocktail party problem", using a single microphone. Our approach involves a deep recurrent neural networks regression to a vector space that is…

声音 · 计算机科学 2017-05-22 Cory Stephenson , Patrick Callier , Abhinav Ganesh , Karl Ni

The ConferencingSpeech 2021 challenge is proposed to stimulate research on far-field multi-channel speech enhancement for video conferencing. The challenge consists of two separate tasks: 1) Task 1 is multi-channel speech enhancement with…

音频与语音处理 · 电气工程与系统科学 2021-04-05 Wei Rao , Yihui Fu , Yanxin Hu , Xin Xu , Yvkai Jv , Jiangyu Han , Zhongjie Jiang , Lei Xie , Yannan Wang , Shinji Watanabe , Zheng-Hua Tan , Hui Bu , Tao Yu , Shidong Shang

Code-switching is a speech phenomenon occurring when a speaker switches language during a conversation. Despite the spontaneous nature of code-switching in conversational spoken language, most existing works collect code-switching data from…

Transitioning between topics is a natural component of human-human dialog. Although topic transition has been studied in dialogue for decades, only a handful of corpora based studies have been performed to investigate the subtleties of…

计算与语言 · 计算机科学 2022-07-21 Mayank Soni , Brendan Spillane , Emer Gilmartin , Christian Saam , Benjamin R. Cowan , Vincent Wade

The Deep Noise Suppression (DNS) challenge is designed to foster innovation in the area of noise suppression to achieve superior perceptual speech quality. We recently organized a DNS challenge special session at INTERSPEECH and ICASSP…

This article introduces the innovative Quantum Dining Information Brokers Problem, presenting a novel entanglement-based quantum protocol to address it. The scenario involves $n$ information brokers, all located in distinct geographical…

For conversational large-vocabulary continuous speech recognition (LVCSR) tasks, up to about two thousand hours of audio is commonly used to train state of the art models. Collection of labeled conversational audio however, is prohibitively…

计算与语言 · 计算机科学 2017-05-30 Shane Walker , Morten Pedersen , Iroro Orife , Jason Flaks

In the field of natural language processing, open-domain chatbots have emerged as an important research topic. However, a major limitation of existing open-domain chatbot research is its singular focus on short single-session dialogue,…

计算与语言 · 计算机科学 2023-10-23 Jihyoung Jang , Minseong Boo , Hyounghun Kim

The success of deep learning has sparked interest in improving relational table tasks, like data preparation and search, with table representation models trained on large table corpora. Existing table corpora primarily contain tables…

数据库 · 计算机科学 2023-04-13 Madelon Hulsebos , Çağatay Demiralp , Paul Groth

We present INDICVOICES, a dataset of natural and spontaneous speech containing a total of 7348 hours of read (9%), extempore (74%) and conversational (17%) audio from 16237 speakers covering 145 Indian districts and 22 languages. Of these…

English is the most widely spoken language in the world, used daily by millions of people as a first or second language in many different contexts. As a result, there are many varieties of English. Although the great many advances in…

计算与语言 · 计算机科学 2023-04-03 Ramon Sanabria , Nikolay Bogoychev , Nina Markl , Andrea Carmantini , Ondrej Klejch , Peter Bell

DeepMine is a speech database in Persian and English designed to build and evaluate text-dependent, text-prompted, and text-independent speaker verification, as well as Persian speech recognition systems. It contains more than 1850 speakers…

音频与语音处理 · 电气工程与系统科学 2019-12-10 Hossein Zeinali , Lukáš Burget , Jan "Honza'' Černocký

Lectures translation is a case of spoken language translation and there is a lack of publicly available parallel corpora for this purpose. To address this, we examine a language independent framework for parallel corpus mining which is a…

计算与语言 · 计算机科学 2020-01-15 Haiyue Song , Raj Dabre , Atsushi Fujita , Sadao Kurohashi

Automatic meeting analysis is an essential fundamental technology required to let, e.g. smart devices follow and respond to our conversations. To achieve an optimal automatic meeting analysis, we previously proposed an all-neural approach…

音频与语音处理 · 电气工程与系统科学 2020-03-10 Keisuke Kinoshita , Marc Delcroix , Shoko Araki , Tomohiro Nakatani