中文
相关论文

相关论文: emg2speech: Synthesizing speech from electromyogra…

200 篇论文

In [1,2] authors provided preliminary results for synthesizing speech from electroencephalography (EEG) features where they first predict acoustic features from EEG features and then the speech is reconstructed from the predicted acoustic…

音频与语音处理 · 电气工程与系统科学 2020-06-03 Gautam Krishna , Co Tran , Mason Carnahan , Ahmed Tewfik

For articulatory-to-acoustic mapping, typically only limited parallel training data is available, making it impossible to apply fully end-to-end solutions like Tacotron2. In this paper, we experimented with transfer learning and adaptation…

音频与语音处理 · 电气工程与系统科学 2021-07-27 Csaba Zainkó , László Tóth , Amin Honarmandi Shandiz , Gábor Gosztolya , Alexandra Markó , Géza Németh , Tamás Gábor Csapó

Self-supervised speech models (S3Ms) are known to encode rich phonetic information, yet how this information is structured remains underexplored. We conduct a comprehensive study across 96 languages to analyze the underlying structure of…

音频与语音处理 · 电气工程与系统科学 2026-04-15 Kwanghee Choi , Eunjung Yeo , Cheol Jun Cho , David Harwath , David R. Mortensen

We present an end-to-end text-to-speech (TTS) synthesis system that generates audio and synchronized tongue motion directly from text. This is achieved by adapting a 3D model of the tongue surface to an articulatory dataset and training a…

人机交互 · 计算机科学 2018-04-17 Ingmar Steiner , Sébastien Le Maguer , Alexander Hewer

Most organisms including humans function by coordinating and integrating sensory signals with motor actions to survive and accomplish desired tasks. Learning these complex sensorimotor mappings proceeds simultaneously and often in an…

音频与语音处理 · 电气工程与系统科学 2023-05-26 Yashish M. Siriwardena , Carol Espy-Wilson , Shihab Shamma

Surface electromyography provides a practical way to infer human movement intention from wearable muscle recordings, but models trained under a single acquisition setting often lose reliability when the user, session, electrode layout, or…

机器学习 · 计算机科学 2026-05-26 Zhenghao Huang , Huilin Yao , Kaikai Wang

Articulatory-to-acoustic (forward) mapping is a technique to predict speech using various articulatory acquisition techniques (e.g. ultrasound tongue imaging, lip video). Real-time MRI (rtMRI) of the vocal tract has not been used before for…

音频与语音处理 · 电气工程与系统科学 2020-08-04 Tamás Gábor Csapó

The electroencephalography (EEG) signals recorded in parallel with speech are used to perform isolated and continuous speech recognition. During speaking process, one also hears his or her own speech and this speech perception is also…

音频与语音处理 · 电气工程与系统科学 2020-06-03 Gautam Krishna , Co Tran , Mason Carnahan , Ahmed Tewfik

Self-supervised speech models (S3Ms) have become an effective backbone for speech applications. Various analyses suggest that S3Ms encode linguistic properties. In this work, we seek a more fine-grained analysis of the word-level linguistic…

计算与语言 · 计算机科学 2024-06-14 Kwanghee Choi , Ankita Pasad , Tomohiko Nakamura , Satoru Fukayama , Karen Livescu , Shinji Watanabe

Covert speech involves imagining speaking without audible sound or any movements. Decoding covert speech from electroencephalogram (EEG) is challenging due to a limited understanding of neural pronunciation mapping and the low…

The relationship between muscle activity and resulting facial expressions is crucial for various fields, including psychology, medicine, and entertainment. The synchronous recording of facial mimicry and muscular activity via surface…

计算机视觉与模式识别 · 计算机科学 2025-03-13 Tim Büchner , Christoph Anders , Orlando Guntinas-Lichius , Joachim Denzler

Silent speech decoding, which performs unvocalized human speech recognition from electroencephalography/electromyography (EEG/EMG), increases accessibility for speech-impaired humans. However, data collection is difficult and performed…

Decoding spoken speech from neural activity in the brain is a fast-emerging research topic, as it could enable communication for people who have difficulties with producing audible speech. For this task, electrocorticography (ECoG) is a…

音频与语音处理 · 电气工程与系统科学 2023-12-22 Miseul Kim , Zhenyu Piao , Jihyun Lee , Hong-Goo Kang

Human infants face a formidable challenge in speech acquisition: mapping extremely variable acoustic inputs into appropriate articulatory movements without explicit instruction. We present a computational model that addresses the…

音频与语音处理 · 电气工程与系统科学 2025-09-16 Marvin Lavechin , Thomas Hueber

Current speech production systems predominantly rely on large transformer models that operate as black boxes, providing little interpretability or grounding in the physical mechanisms of human speech. We address this limitation by proposing…

音频与语音处理 · 电气工程与系统科学 2025-10-08 Akshay Anand , Chenxu Guo , Cheol Jun Cho , Jiachen Lian , Gopala Anumanchipalli

Recently, speech representation learning has improved many speech-related tasks such as speech recognition, speech classification, and speech-to-text translation. However, all the above tasks are in the direction of speech understanding,…

音频与语音处理 · 电气工程与系统科学 2022-06-22 He Bai , Renjie Zheng , Junkun Chen , Xintong Li , Mingbo Ma , Liang Huang

Audio Large Language Models (Audio LLMs) have demonstrated strong capabilities in integrating speech perception with language understanding. However, whether their internal representations align with human neural dynamics during…

声音 · 计算机科学 2026-02-04 Haoyun Yang , Xin Xiao , Jiang Zhong , Yu Tian , Dong Xiaohua , Yu Mao , Hao Wu , Kaiwen Wei

Understanding the underlying relationship between tongue and oropharyngeal muscle deformation seen in tagged-MRI and intelligible speech plays an important role in advancing speech motor control theories and treatment of speech…

Speech decoding from EEG signals is a challenging task, where brain activity is modeled to estimate salient characteristics of acoustic stimuli. We propose FESDE, a novel framework for Fully-End-to-end Speech Decoding from EEG signals. Our…

信号处理 · 电气工程与系统科学 2024-06-14 Jihwan Lee , Aditya Kommineni , Tiantian Feng , Kleanthis Avramidis , Xuan Shi , Sudarsana Kadiri , Shrikanth Narayanan

We propose a computational model of speech production combining a pre-trained neural articulatory synthesizer able to reproduce complex speech stimuli from a limited set of interpretable articulatory parameters, a DNN-based internal forward…

声音 · 计算机科学 2022-04-06 Marc-Antoine Georges , Julien Diard , Laurent Girin , Jean-Luc Schwartz , Thomas Hueber