English
Related papers

Related papers: Synthesizing Speech from Intracranial Depth Electr…

200 papers

Decoding neural activity into human-interpretable representations is a key research direction in brain-computer interfaces (BCIs) and computational neuroscience. Recent progress in machine learning and generative AI has driven growing…

Artificial Intelligence · Computer Science 2025-12-02 Shreya Shukla , Jose Torres , Akshaj Murhekar , Christina Liu , Abhijit Mishra , Jacek Gwizdka , Shounak Roychowdhury

The rapid spread of media content synthesis technology and the potentially damaging impact of audio and video deepfakes on people's lives have raised the need to implement systems able to detect these forgeries automatically. In this work…

Sound · Computer Science 2022-11-01 Luigi Attorresi , Davide Salvi , Clara Borrelli , Paolo Bestagini , Stefano Tubaro

Intracranial electrocorticography (ECoG) offers high-signal-to-noise access to cortical activity for brain-computer interfaces, yet limited per-patient data has led most prior work to rely on small, subject-specific decoders that neglect…

Artificial Intelligence · Computer Science 2026-05-12 Liuyin Yang , Qiang Sun , Bob Van Dyck , Eva Calvo Merino , Marc M. Van Hulle

Decoding the directional focus of an attended speaker from listeners' electroencephalogram (EEG) signals is essential for developing brain-computer interfaces to improve the quality of life for individuals with hearing impairment. Previous…

Sound · Computer Science 2025-10-23 Yuanming Zhang , Jing Lu , Fei Chen , Haoliang Du , Xia Gao , Zhibin Lin

Electroencephalography (EEG) is a neuroimaging technique that records brain neural activity with high temporal resolution. Unlike other methods, EEG does not require prohibitively expensive equipment and can be easily set up using…

Human-Computer Interaction · Computer Science 2024-10-01 Arash Akbarinia

The amount of articulatory data available for training deep learning models is much less compared to acoustic speech data. In order to improve articulatory-to-acoustic synthesis performance in these low-resource settings, we propose a…

Audio and Speech Processing · Electrical Eng. & Systems 2024-12-19 Peter Wu , Bohan Yu , Kevin Scheck , Alan W Black , Aditi S. Krishnapriyan , Irene Y. Chen , Tanja Schultz , Shinji Watanabe , Gopala K. Anumanchipalli

The conversion from text to speech relies on the accurate mapping from linguistic to acoustic symbol sequences, for which current practice employs recurrent statistical models like recurrent neural networks. Despite the good performance of…

Sound · Computer Science 2018-11-07 Santiago Pascual , Antonio Bonafonte , Joan Serrà

This paper proposes a speech rhythm-based method for speaker embeddings to model phoneme duration using a few utterances by the target speaker. Speech rhythm is one of the essential factors among speaker characteristics, along with acoustic…

Sound · Computer Science 2024-02-13 Kenichi Fujita , Atsushi Ando , Yusuke Ijima

Discrete speech representation learning has recently attracted increasing interest in both acoustic and semantic modeling. Existing approaches typically encode 16 kHz waveforms into discrete tokens at a rate of 25 or 50 tokens per second.…

Computation and Language · Computer Science 2025-09-03 Jialong Zuo , Guangyan Zhang , Minghui Fang , Shengpeng Ji , Xiaoqi Jiao , Jingyu Li , Yiwen Guo , Zhou Zhao

As an indispensable part of modern human-computer interaction system, speech synthesis technology helps users get the output of intelligent machine more easily and intuitively, thus has attracted more and more attention. Due to the…

Sound · Computer Science 2021-04-21 Zhaoxi Mu , Xinyu Yang , Yizhuo Dong

The electroencephalography (EEG) signals recorded in parallel with speech are used to perform isolated and continuous speech recognition. During speaking process, one also hears his or her own speech and this speech perception is also…

Audio and Speech Processing · Electrical Eng. & Systems 2020-06-03 Gautam Krishna , Co Tran , Mason Carnahan , Ahmed Tewfik

Brain Computer Interface (BCI) can help patients of neuromuscular diseases restore parts of the movement and communication abilities that they have lost. Most of BCIs rely on mapping brain activities to device instructions, but limited…

Human-Computer Interaction · Computer Science 2017-05-23 Kang Wang , Xueqian Wang , Gang Li

Decoding visual information from electroencephalography (EEG) signals remains a fundamental challenge in brain-computer interfaces and medical rehabilitation. Existing EEG visual decoding methods mainly focus on learning a single global EEG…

Computer Vision and Pattern Recognition · Computer Science 2026-05-22 Xiang Gao , Hui Tian , Yanming Zhu , Xuefei Yin , Alan Wee-Chung Liew

To investigate how speech is processed in the brain, we can model the relation between features of a natural speech signal and the corresponding recorded electroencephalogram (EEG). Usually, linear models are used in regression tasks.…

Audio and Speech Processing · Electrical Eng. & Systems 2022-07-25 Corentin Puffay , Jana Van Canneyt , Jonas Vanthornhout , Hugo Van Hamme , Tom Francart

Surface electromyography (EMG) is a promising modality for silent speech interfaces, but its effectiveness depends heavily on sensor placement and channel availability. In this work, we investigate the contribution of individual and…

Sound · Computer Science 2026-02-09 Injune Hwang , Jaejun Lee , Kyogu Lee

High-quality speech corpora are essential foundations for most speech applications. However, such speech data are expensive and limited since they are collected in professional recording environments. In this work, we propose an…

Audio and Speech Processing · Electrical Eng. & Systems 2020-11-11 Haoyu Li , Yang Ai , Junichi Yamagishi

OBJECTIVE: We aim to extract and denoise the attended speaker in a noisy, two-speaker acoustic scenario, relying on microphone array recordings from a binaural hearing aid, which are complemented with electroencephalography (EEG) recordings…

Sound · Computer Science 2019-02-06 Simon Van Eyndhoven , Tom Francart , Alexander Bertrand

Deep learning based models have significantly improved the performance of speech separation with input mixtures like the cocktail party. Prominent methods (e.g., frequency-domain and time-domain speech separation) usually build regression…

Sound · Computer Science 2022-01-11 Jing Shi , Xuankai Chang , Tomoki Hayashi , Yen-Ju Lu , Shinji Watanabe , Bo Xu

Silent Speech Decoding (SSD), based on articulatory neuromuscular activities, has become a prevalent task of Brain-Computer Interface (BCI) in recent years. Many works have been devoted to decoding surface electromyography (sEMG) from…

Sound · Computer Science 2022-06-02 Huiyan Li , Haohong Lin , You Wang , Hengyang Wang , Ming Zhang , Han Gao , Qing Ai , Zhiyuan Luo , Guang Li

Objective: Currently, only behavioral speech understanding tests are available, which require active participation of the person being tested. As this is infeasible for certain populations, an objective measure of speech intelligibility is…

Audio and Speech Processing · Electrical Eng. & Systems 2021-11-30 Bernd Accou , Mohammad Jalilpour Monesi , Hugo Van hamme , Tom Francart