English
Related papers

Related papers: Cortical Temporal Mismatch Compensation in Bimodal…

200 papers

Although many previous studies have carried out multimodal learning with real-time MRI data that captures the audio-visual kinematics of the vocal tract during speech, these studies have been limited by their reliance on multi-speaker…

Machine Learning · Computer Science 2025-06-02 Sean Foley , Hong Nguyen , Jihwan Lee , Sudarsana Reddy Kadiri , Dani Byrd , Louis Goldstein , Shrikanth Narayanan

Spoofing countermeasure (CM) systems are critical in speaker verification; they aim to discern spoofing attacks from bona fide speech trials. In practice, however, acoustic condition variability in speech utterances may significantly…

Audio and Speech Processing · Electrical Eng. & Systems 2026-02-05 You Zhang , Ge Zhu , Fei Jiang , Zhiyao Duan

Speaker verification has been widely explored using speech signals, which has shown significant improvement using deep models. Recently, there has been a surge in exploring faces and voices as they can offer more complementary and…

Sound · Computer Science 2023-09-29 R. Gnana Praveen , Jahangir Alam

Steady-State Visual Evoked Potential (SSVEP) spellers are a promising communication tool for individuals with disabilities. This Brain-Computer Interface utilizes scalp potential data from (electroencephalography) EEG electrodes on a…

Human-Computer Interaction · Computer Science 2024-12-31 Joseph Zhang , Ruiming Zhang , Kipngeno Koech , David Hill , Kateryna Shapovalenko

As synchronized activity is associated with basic brain functions and pathological states, spike train synchrony has become an important measure to analyze experimental neuronal data. Many different measures of spike train synchrony have…

The brain exhibits capabilities of fast incremental learning from few noisy examples, as well as the ability to associate similar memories in autonomously-created categories and to combine contextual hints with sensory perceptions. Together…

Trouble hearing in noisy situations remains a common complaint for both individuals with hearing loss and individuals with normal hearing. This is hypothesized to arise due to condition called: cochlear neural degeneration (CND) which can…

Audio and Speech Processing · Electrical Eng. & Systems 2025-06-17 Ahsan J. Cheema , Sunil Puria

Multi-channel audio alignment is a key requirement in bioacoustic monitoring, spatial audio systems, and acoustic localization. However, existing methods often struggle to address nonlinear clock drift and lack mechanisms for quantifying…

Sound · Computer Science 2025-09-23 Ragib Amin Nihal , Benjamin Yen , Takeshi Ashizawa , Kazuhiro Nakadai

Speech intelligibility can be affected by multiple factors, such as noisy environments, channel distortions or physiological issues. In this work, we deal with the problem of automatic prediction of the speech intelligibility level in this…

Audio and Speech Processing · Electrical Eng. & Systems 2024-02-06 Ascensión Gallardo-Antolín , Juan M. Montero

Audio-visual feature synchronization for real-time speech enhancement in hearing aids represents a progressive approach to improving speech intelligibility and user experience, particularly in strong noisy backgrounds. This approach…

Audio and Speech Processing · Electrical Eng. & Systems 2025-08-28 Nasir Saleem , Mandar Gogate , Kia Dashtipour , Adeel Hussain , Usman Anwar , Adewale Adetomi , Tughrul Arslan , Amir Hussain

Multi-resolution spectro-temporal features of a speech signal represent how the brain perceives sounds by tuning cortical cells to different spectral and temporal modulations. These features produce a higher dimensional representation of…

Audio and Speech Processing · Electrical Eng. & Systems 2022-06-28 Rahil Parikh , Nadee Seneviratne , Ganesh Sivaraman , Shihab Shamma , Carol Espy-Wilson

We introduce a multiple criteria Bayesian preference learning framework incorporating behavioral cues for decision aiding. The framework integrates pairwise comparisons, response time, and attention duration to deepen insights into…

Applications · Statistics 2025-04-22 Jiaxuan Jiang , Jiapeng Liu , Miłosz Kadziński , Xiuwu Liao , Jingyu Dong

Designing a feedback that helps participants to achieve higher performances is an important concern in brain-computer interface (BCI) research. In a pilot study, we demonstrate how a congruent auditory feedback could improve classification…

To improve speech intelligibility in complex everyday situations, the human auditory system partially relies on Interaural Time Differences (ITDs) and Interaural Level Differences (ILDs). However, hearing impaired (HI) listeners often…

Audio and Speech Processing · Electrical Eng. & Systems 2025-11-27 Timm-Jonas Bäumer , Johannes W. de Vries , Stephan Töpken , Richard C. Hendriks , Peyman Goli , Steven van de Par

In this paper, we aim to unveil the impact of data augmentation in audio-language multi-modal learning, which has not been explored despite its importance. We explore various augmentation methods at not only train-time but also test-time…

Sound · Computer Science 2023-05-24 Eungbeom Kim , Jinhee Kim , Yoori Oh , Kyungsu Kim , Minju Park , Jaeheon Sim , Jinwoo Lee , Kyogu Lee

Recent progress has been made in detecting early stage dementia entirely through recordings of patient speech. Multimodal speech analysis methods were applied to the PROCESS challenge, which requires participants to use audio recordings of…

Audio and Speech Processing · Electrical Eng. & Systems 2025-02-14 Lei Chi , Arav Sharma , Ari Gebhardt , Joseph T. Colonel

Detecting collaborative and problem-solving behaviours from digital traces to interpret students' collaborative problem solving (CPS) competency is a long-term goal in the Artificial Intelligence in Education (AIEd) field. Although…

Computation and Language · Computer Science 2025-04-22 K. Wong , B. Wu , S. Bulathwela , M. Cukurova

Circadian and other physiological rhythms play a key role in both normal homeostasis and disease processes. Such is the case of circadian and infradian seizure patterns observed in epilepsy. In this paper we explore a new implantable…

Given an unknown audio source, the estimation of time differences-of-arrivals (TDOAs) can be efficiently and robustly solved using blind channel identification and exploiting the cross-correlation identity (CCI). Prior "blind" works have…

Sound · Computer Science 2020-10-19 Danilo Greco , Jacopo Cavazza , Alessio Del Bue

We introduce Purrfect Pitch, a system consisting of a wearable haptic device and a custom-designed learning interface for musical ear training. We focus on the ability to identify musical intervals (sequences of two musical notes), which is…

Human-Computer Interaction · Computer Science 2024-07-16 Sam Chin , Cathy Mengying Fang , Nikhil Singh , Ibrahim Ibrahim , Joe Paradiso , Pattie Maes