English
Related papers

Related papers: PROCESS-2: A Benchmark Speech Corpus for Early Cog…

200 papers

This paper presents the development and validation of an eye-tracking dataset designed to investigate how second-language (L2) learners process idiomatic expressions. While native speakers often rely on direct retrieval of figurative…

Computation and Language · Computer Science 2026-05-07 Eduardo Santos , Juliana Carvalho , César Rennó-Costa

Speech recognition is very challenging in student learning environments that are characterized by significant cross-talk and background noise. To address this problem, we present a bilingual speech recognition system that uses an…

Speech foundation models have demonstrated exceptional capabilities in speech-related tasks. Nevertheless, these models often struggle with non-verbal audio data, such as vocalizations, baby crying, etc., which are critical for various…

Audio and Speech Processing · Electrical Eng. & Systems 2025-02-25 Alkis Koudounas , Moreno La Quatra , Marco Sabato Siniscalchi , Elena Baralis

Access to informative databases is a crucial part of notable research developments. In the field of domestic audio classification, there have been significant advances in recent years. Although several audio databases exist, these can be…

Audio and Speech Processing · Electrical Eng. & Systems 2021-05-14 Abigail Copiaco , Christian Ritz , Stefano Fasciani , Nidhal Abdulaziz

We target passive dementia screening from short camera-facing talking head video, developing a facial temporal micro dynamics analysis for language free detection of early neuro cognitive change. This enables unscripted, in the wild video…

Computer Vision and Pattern Recognition · Computer Science 2025-11-27 Filippo Cenacchi , Longbing Cao , Mitchell McEwan , Deborah Richards

Speech emotion recognition is a vital contributor to the next generation of human-computer interaction (HCI). However, current existing small-scale databases have limited the development of related research. In this paper, we present LSSED,…

Sound · Computer Science 2021-02-04 Weiquan Fan , Xiangmin Xu , Xiaofen Xing , Weidong Chen , Dongyan Huang

In the past decade, there has been a surge in research examining the use of voice and speech analysis as a means of detecting neurodegenerative diseases such as Alzheimer's. Many studies have shown that certain acoustic features can be used…

Audio and Speech Processing · Electrical Eng. & Systems 2023-07-20 Vrindha M. K. , Geethu V. , Anurenjan P. R. , Deepak S. , Sreeni K. G.

Lip-to-speech (L2S) synthesis, which reconstructs speech from visual cues, faces challenges in accuracy and naturalness due to limited supervision in capturing linguistic content, accents, and prosody. In this paper, we propose RESOUND, a…

Sound · Computer Science 2025-05-29 Long-Khanh Pham , Thanh V. T. Tran , Minh-Tan Pham , Van Nguyen

Decoding speech from brain activity is a long-awaited goal in both healthcare and neuroscience. Invasive devices have recently led to major milestones in that regard: deep learning algorithms trained on intracranial recordings now start to…

Audio and Speech Processing · Electrical Eng. & Systems 2023-10-06 Alexandre Défossez , Charlotte Caucheteux , Jérémy Rapin , Ori Kabeli , Jean-Rémi King

Speech translation has recently become an increasingly popular topic of research, partly due to the development of benchmark datasets. Nevertheless, current datasets cover a limited number of languages. With the aim to foster research in…

Computation and Language · Computer Science 2020-10-27 Changhan Wang , Anne Wu , Juan Pino

Decoding brain activity into natural language is a major challenge in AI with important applications in assistive communication, neurotechnology, and human-computer interaction. Most existing Brain-Computer Interface (BCI) approaches rely…

Machine Learning · Computer Science 2026-03-19 Akshaj Murhekar , Christina Liu , Abhijit Mishra , Shounak Roychowdhury , Jacek Gwizdka

We present the LEMAS-Dataset, which, to our knowledge, is currently the largest open-source multilingual speech corpus with word-level timestamps. Covering over 150,000 hours across 10 major languages, LEMAS-Dataset is constructed via a…

Sound · Computer Science 2026-01-09 Zhiyuan Zhao , Lijian Lin , Ye Zhu , Kai Xie , Yunfei Liu , Yu Li

People spend a substantial portion of their lives engaged in conversation, and yet our scientific understanding of conversation is still in its infancy. In this report we advance an interdisciplinary science of conversation, with findings…

Computation and Language · Computer Science 2022-03-02 Andrew Reece , Gus Cooney , Peter Bull , Christine Chung , Bryn Dawson , Casey Fitzpatrick , Tamara Glazer , Dean Knox , Alex Liebscher , Sebastian Marin

Emergency Medical Services (EMS) responders often operate under time-sensitive conditions, facing cognitive overload and inherent risks, requiring essential skills in critical thinking and rapid decision-making. This paper presents…

Artificial Intelligence · Computer Science 2024-10-27 Keshara Weerasinghe , Saahith Janapati , Xueren Ge , Sion Kim , Sneha Iyer , John A. Stankovic , Homa Alemzadeh

Standardized tests play a crucial role in the detection of cognitive impairment. Previous work demonstrated that automatic detection of cognitive impairment is possible using audio data from a standardized picture description task. The…

Training data cleaning is a new application for generative model-based speech restoration (SR). This paper introduces Miipher-2, an SR model designed for million-hour scale data, for training data cleaning for large-scale generative models…

The success of large language models has driven interest in developing similar speech processing capabilities. However, a key challenge is the scarcity of high-quality spontaneous speech data, as most existing datasets contain scripted…

Vision-guided speech generation aims to produce authentic speech from facial appearance or lip motions without relying on auditory signals, offering significant potential for applications such as dubbing in filmmaking and assisting…

Computer Vision and Pattern Recognition · Computer Science 2025-03-20 Jiaxin Ye , Hongming Shan

The increasing adoption of text-to-speech technologies has led to a growing demand for natural and emotive voices that adapt to a conversation's context and emotional tone. The Emotive Narrative Storytelling (EMNS) corpus is a unique speech…

Computation and Language · Computer Science 2023-05-26 Kari Ali Noriy , Xiaosong Yang , Jian Jun Zhang

Excessive sleepiness in attention-critical contexts can lead to adverse events, such as car crashes. Detecting and monitoring sleepiness can help prevent these adverse events from happening. In this paper, we use the Voiceome dataset to…

Computation and Language · Computer Science 2021-11-30 Bang Tran , Youxiang Zhu , Xiaohui Liang , James W. Schwoebel , Lindsay A. Warrenburg