中文
相关论文

相关论文: DiPCo -- Dinner Party Corpus

200 篇论文

In this paper, we present a transcribed corpus of the LIBE committee of the EU parliament, totalling 3.6 Million running words. The meetings of parliamentary committees of the EU are a potentially valuable source of information for…

计算与语言 · 计算机科学 2023-04-18 Hugo de Vos , Suzan Verberne

In this work, we present a new dataset for conversational recommendation over knowledge graphs in e-commerce platforms called COOKIE. The dataset is constructed from an Amazon review corpus by integrating both user-agent dialogue and custom…

信息检索 · 计算机科学 2020-08-24 Zuohui Fu , Yikun Xian , Yaxin Zhu , Yongfeng Zhang , Gerard de Melo

Human conversations are complicated and building a human-like dialogue agent is an extremely challenging task. With the rapid development of deep learning techniques, data-driven models become more and more prevalent which need a huge…

计算与语言 · 计算机科学 2020-03-25 Meng Chen , Ruixue Liu , Lei Shen , Shaozu Yuan , Jingyan Zhou , Youzheng Wu , Xiaodong He , Bowen Zhou

It is now a common practice to compare models of human language processing by predicting participant reactions (such as reading times) to corpora consisting of rich naturalistic linguistic materials. However, many of the corpora used in…

This paper describes the Spot the Difference Corpus which contains 54 interactions between pairs of subjects interacting to find differences in two very similar scenes. The setup used, the participants' metadata and details about collection…

计算与语言 · 计算机科学 2018-05-15 José Lopes , Nils Hemmingsson , Oliver Åstrand

Since its introduction in 2018, EPIC-KITCHENS has attracted attention as the largest egocentric video benchmark, offering a unique viewpoint on people's interaction with objects, their attention, and even intention. In this paper, we detail…

Recent progress in speech processing has highlighted that high-quality performance across languages requires substantial training data for each individual language. While existing multilingual datasets cover many languages, they often…

计算与语言 · 计算机科学 2025-10-28 Samuel Pfisterer , Florian Grötschla , Luca A. Lanzendörfer , Florian Yan , Roger Wattenhofer

We describe a private audio messaging system that uses echoes to unscramble messages at a few predetermined locations in a room. The system works by splitting the audio into short chunks and emitting them from different loudspeakers. The…

声音 · 计算机科学 2018-09-18 Yu-Jeh Liu , Jonah Casebeer , Ivan Dokmanić

In this paper, we propose the scheme for annotating large-scale multi-party chat dialogues for discourse parsing and machine comprehension. The main goal of this project is to help understand multi-party dialogues. Our dataset is based on…

计算与语言 · 计算机科学 2019-11-12 Jiaqi Li , Ming Liu , Bing Qin , Zihao Zheng , Ting Liu

Conversational Product Search ( CPS ) systems interact with users via natural language to offer personalized and context-aware product lists. However, most existing research on CPS is limited to simulated conversations, due to the lack of a…

计算与语言 · 计算机科学 2025-04-29 Jie Zou , Mohammad Aliannejadi , Evangelos Kanoulas , Shuxi Han , Heli Ma , Zheng Wang , Yang Yang , Heng Tao Shen

Disentangling conversations mixed together in a single stream of messages is a difficult task, made harder by the lack of large manually annotated datasets. We created a new dataset of 77,563 messages manually annotated with reply-structure…

In this paper, we present AISHELL-4, a sizable real-recorded Mandarin speech dataset collected by 8-channel circular microphone array for speech processing in conference scenario. The dataset consists of 211 recorded meeting sessions, each…

声音 · 计算机科学 2021-08-11 Yihui Fu , Luyao Cheng , Shubo Lv , Yukai Jv , Yuxiang Kong , Zhuo Chen , Yanxin Hu , Lei Xie , Jian Wu , Hui Bu , Xin Xu , Jun Du , Jingdong Chen

This paper presents an extension to a very low-resource parallel corpus collected in an endangered language, Griko, making it useful for computational research. The corpus consists of 330 utterances (about 20 minutes of speech) which have…

Mental distress like depression and anxiety contribute to the largest proportion of the global burden of diseases. Automated diagnosis systems of such disorders, empowered by recent innovations in Artificial Intelligence, can pave the way…

音频与语音处理 · 电气工程与系统科学 2023-06-23 Mashrura Tasnim , Malikeh Ehghaghi , Brian Diep , Jekaterina Novikova

The "MEG-MASC" dataset provides a curated set of raw magnetoencephalography (MEG) recordings of 27 English speakers who listened to two hours of naturalistic stories. Each participant performed two identical sessions, involving listening to…

定量方法 · 定量生物学 2022-08-25 Laura Gwilliams , Graham Flick , Alec Marantz , Liina Pylkkanen , David Poeppel , Jean-Remi King

Dubbed series are gaining a lot of popularity in recent years with strong support from major media service providers. Such popularity is fueled by studies that showed that dubbed versions of TV shows are more popular than their subtitled…

计算与语言 · 计算机科学 2022-03-08 Massa Baali , Wassim El-Hajj , Ahmed Ali

Dementia affects cognitive functions of adults, including memory, language, and behaviour. Standard diagnostic biomarkers such as MRI are costly, whilst neuropsychological tests suffer from sensitivity issues in detecting dementia onset.…

计算与语言 · 计算机科学 2023-12-27 Dimitris Gkoumas , Bo Wang , Adam Tsakalidis , Maria Wolters , Arkaitz Zubiaga , Matthew Purver , Maria Liakata

This article describes the MyST corpus developed as part of the My Science Tutor project -- one of the largest collections of children's conversational speech comprising approximately 400 hours, spanning some 230K utterances across about…

计算与语言 · 计算机科学 2023-09-26 Sameer S. Pradhan , Ronald A. Cole , Wayne H. Ward

We present the Latvian Twitter Eater Corpus - a set of tweets in the narrow domain related to food, drinks, eating and drinking. The corpus has been collected over time-span of over 8 years and includes over 2 million tweets entailed with…

计算与语言 · 计算机科学 2020-09-02 Uga Sproģis , Matīss Rikters

The scarcity of large-scale classroom speech data has hindered the development of AI-driven speech models for education. Public classroom datasets remain limited, and the lack of a dedicated classroom noise corpus prevents the use of…

声音 · 计算机科学 2025-06-12 Ahmed Adel Attia , Jing Liu , Carl Espy-Wilson