English
Related papers

Related papers: Relating the fundamental frequency of speech with …

200 papers

Neuroscientists have recently turned to intracranial brain recording methods, like electrocorticography (ECoG), for human experiments because of the fine spatial and temporal resolution that they afford. Models trained on this data,…

Computation and Language · Computer Science 2026-05-20 Aditya R. Vaidya , Richard J. Antonello , Alexander G. Huth

Speech emotion recognition is an important and challenging task in the realm of human-computer interaction. Prior work proposed a variety of models and feature sets for training a system. In this work, we conduct extensive experiments using…

Computation and Language · Computer Science 2017-06-05 Michael Neumann , Ngoc Thang Vu

Epilepsy is one of the most common neurological disorders. This disease requires reliable and efficient seizure detection methods. Electroencephalography (EEG) is the gold standard for seizure monitoring, but its manual analysis is a…

Signal Processing · Electrical Eng. & Systems 2025-12-17 Annika Stiehl , Nicolas Weeger , Christian Uhl , Dominic Bechtold , Nicole Ille , Stefan Geißelsöder

In this review, we examine computational models that explore the role of neural oscillations in speech perception, spanning from early auditory processing to higher cognitive stages. We focus on models that use rhythmic brain activities,…

Neurons and Cognition · Quantitative Biology 2025-02-19 Olesia Dogonasheva , Denis Zakharov , Anne-Lise Giraud , Boris Gutkin

Speech foundation models (SFMs) are increasingly hailed as powerful computational models of human speech perception. However, since their representations are inherently black-box, it remains unclear what drives their alignment with brain…

Neurons and Cognition · Quantitative Biology 2025-09-26 Riki Shimizu , Richard J. Antonello , Chandan Singh , Nima Mesgarani

High-quality speech corpora are essential foundations for most speech applications. However, such speech data are expensive and limited since they are collected in professional recording environments. In this work, we propose an…

Audio and Speech Processing · Electrical Eng. & Systems 2020-11-11 Haoyu Li , Yang Ai , Junichi Yamagishi

Electroencephalography (EEG) is one of the most common signals used to capture the electrical activity of the brain, and the decoding of EEG, to acquire the user intents, has been at the forefront of brain-computer/machine interfaces…

Machine Learning · Computer Science 2025-07-04 Haodong Zhang , Hongqi Li

Decoding inner speech from the brain signal via hybridisation of fMRI and EEG data is explored to investigate the performance benefits over unimodal models. Two different bimodal fusion approaches are examined: concatenation of probability…

Invasive brain-computer interfaces with Electrocorticography (ECoG) have shown promise for high-performance speech decoding in medical applications, but less damaging methods like intracranial stereo-electroencephalography (sEEG) remain…

Signal Processing · Electrical Eng. & Systems 2024-11-04 Hui Zheng , Hai-Teng Wang , Wei-Bang Jiang , Zhong-Tao Chen , Li He , Pei-Yang Lin , Peng-Hu Wei , Guo-Guang Zhao , Yun-Zhe Liu

Brain-computer interfaces (BCIs) hold great potential for aiding individuals with speech impairments. Utilizing electroencephalography (EEG) to decode speech is particularly promising due to its non-invasive nature. However, recordings are…

Neurons and Cognition · Quantitative Biology 2024-07-11 Motoshige Sato , Kenichi Tomeoka , Ilya Horiguchi , Kai Arulkumaran , Ryota Kanai , Shuntaro Sasai

Electroencephalography (EEG) signals reflect activities on certain brain areas. Effective classification of time-varying EEG signals is still challenging. First, EEG signal processing and feature engineering are time-consuming and highly…

Human-Computer Interaction · Computer Science 2019-08-27 Xiang Zhang , Lina Yao , Xianzhi Wang , Wenjie Zhang , Shuai Zhang , Yunhao Liu

Talking head generation is a significant research topic that still faces numerous challenges. Previous works often adopt generative adversarial networks or regression models, which are plagued by generation quality and average facial shape…

Computer Vision and Pattern Recognition · Computer Science 2024-08-20 Ziyu Yao , Xuxin Cheng , Zhiqi Huang

Accurate prediction of epileptic seizures has remained elusive, despite the many advances in machine learning and time-series classification. In this work, we develop a convolutional network module that exploits Electroencephalogram (EEG)…

Image and Video Processing · Electrical Eng. & Systems 2020-07-24 Ramy Hussein , Soojin Lee , Rabab Ward , Martin J. McKeown

We explore whether neural networks can decode brain activity into speech by mapping EEG recordings to audio representations. Using EEG data recorded as subjects listened to natural speech, we train a model with a contrastive CLIP loss to…

Sound · Computer Science 2025-11-10 Quentin Auster , Kateryna Shapovalenko , Chuang Ma , Demaio Sun

Deep learning for decoding EEG signals has gained traction, with many claims to state-of-the-art accuracy. However, despite the convincing benchmark performance, successful translation to real applications is limited. The frequent…

Premise. Patterns of electrical brain activity recorded via electroencephalography (EEG) offer immense value for scientific and clinical investigations. The inability of supervised EEG encoders to learn robust EEG patterns and their…

Signal Processing · Electrical Eng. & Systems 2025-12-25 Gayal Kuruppu , Neeraj Wagh , Vaclav Kremen , Sandipan Pati , Gregory Worrell , Yogatheesan Varatharajah

Recent speech-to-speech (S2S) models generate intelligible speech but still lack natural expressiveness, largely due to the absence of a reliable evaluation metric. Existing approaches, such as subjective MOS ratings, low-level acoustic…

Sound · Computer Science 2025-10-24 Zhiyu Lin , Jingwen Yang , Jiale Zhao , Meng Liu , Sunzhu Li , Benyou Wang

With recent research advancements, deep learning models are becoming attractive and powerful choices for speech enhancement in real-time applications. While state-of-the-art models can achieve outstanding results in terms of speech quality…

Audio and Speech Processing · Electrical Eng. & Systems 2021-05-20 Sebastian Braun , Hannes Gamper , Chandan K. A. Reddy , Ivan Tashev

Sound event detection systems typically consist of two stages: extracting hand-crafted features from the raw audio waveform, and learning a mapping between these features and the target sound events using a classifier. Recently, the focus…

Sound · Computer Science 2018-05-11 Emre Çakır , Tuomas Virtanen

Automatic speech recognition in reverberant conditions is a challenging task as the long-term envelopes of the reverberant speech are temporally smeared. In this paper, we propose a neural model for enhancement of sub-band temporal…

Audio and Speech Processing · Electrical Eng. & Systems 2020-08-11 Anurenjan Purushothaman , Anirudh Sreeram , Rohit Kumar , Sriram Ganapathy