English
Related papers

Related papers: emg2speech: Synthesizing speech from electromyogra…

200 papers

Articulatory-to-acoustic mapping seeks to reconstruct speech from a recording of the articulatory movements, for example, an ultrasound video. Just like speech signals, these recordings represent not only the linguistic content, but are…

We present a novel approach for synthesizing 3D facial motions from audio sequences using key motion embeddings. Despite recent advancements in data-driven techniques, accurately mapping between audio signals and 3D facial meshes remains…

Computer Vision and Pattern Recognition · Computer Science 2024-09-04 Zhihao Xu , Shengjie Gong , Jiapeng Tang , Lingyu Liang , Yining Huang , Haojie Li , Shuangping Huang

Articulatory acoustic inversion aims to reconstruct the complete geometry of the vocal tract from the speech signal. In this paper, we present a comparative study of several levels of phonetic segmentation accuracy, together with a…

Audio and Speech Processing · Electrical Eng. & Systems 2026-03-13 Sofiane Azzouz , Pierre-André Vuissoz , Yves Laprie

Emotions play a central role in human communication, shaping trust, engagement, and social interaction. As artificial intelligence systems powered by large language models become increasingly integrated into everyday life, enabling them to…

Audio and Speech Processing · Electrical Eng. & Systems 2026-03-11 Soumya Dutta

Spoken language models (SLMs) that integrate speech with large language models (LMs) rely on modality adapters (MAs) to map the output of speech encoders to a representation that is understandable to the decoder LM. Yet we know very little…

Computation and Language · Computer Science 2025-10-20 Tolúlopé Ògúnrèmí , Christopher D. Manning , Dan Jurafsky , Karen Livescu

Attention-based encoder-decoder model has achieved impressive results for both automatic speech recognition (ASR) and text-to-speech (TTS) tasks. This approach takes advantage of the memorization capacity of neural networks to learn the…

Computation and Language · Computer Science 2020-03-17 Chengyi Wang , Yu Wu , Yujiao Du , Jinyu Li , Shujie Liu , Liang Lu , Shuo Ren , Guoli Ye , Sheng Zhao , Ming Zhou

This paper introduces a novel cross-physiology translation task: synthesizing sleep electroencephalography (EEG) from respiration signals. To address the significant complexity gap between the two modalities, we propose a…

Machine Learning · Computer Science 2026-02-03 Kaiwen Zha , Chao Li , Hao He , Peng Cao , Tianhong Li , Ali Mirzazadeh , Ellen Zhang , Jong Woo Lee , Yoon Kim , Dina Katabi

We describe a neural network-based system for text-to-speech (TTS) synthesis that is able to generate speech audio in the voice of many different speakers, including those unseen during training. Our system consists of three independently…

Computation and Language · Computer Science 2019-01-04 Ye Jia , Yu Zhang , Ron J. Weiss , Quan Wang , Jonathan Shen , Fei Ren , Zhifeng Chen , Patrick Nguyen , Ruoming Pang , Ignacio Lopez Moreno , Yonghui Wu

Vocal tract articulation is a natural, grounded control space of speech production. The spatiotemporal coordination of articulators combined with the vocal source shapes intelligible speech sounds to enable effective spoken communication.…

Audio and Speech Processing · Electrical Eng. & Systems 2025-03-04 Cheol Jun Cho , Peter Wu , Tejas S. Prabhune , Dhruv Agarwal , Gopala K. Anumanchipalli

In this paper, we propose a novel Lip-to-Speech synthesis (L2S) framework, for synthesizing intelligible speech from a silent lip movement video. Specifically, to complement the insufficient supervisory signal of the previous L2S model, we…

Sound · Computer Science 2023-06-01 Jeongsoo Choi , Minsu Kim , Yong Man Ro

This paper presents Translatotron 3, a novel approach to unsupervised direct speech-to-speech translation from monolingual speech-text datasets by combining masked autoencoder, unsupervised embedding mapping, and back-translation.…

Computation and Language · Computer Science 2024-01-17 Eliya Nachmani , Alon Levkovitch , Yifan Ding , Chulayuth Asawaroengchai , Heiga Zen , Michelle Tadmor Ramanovich

Self-supervised approaches for electroencephalography (EEG) representation learning face three specific challenges inherent to EEG data: (1) The low signal-to-noise ratio which challenges the quality of the representation learned, (2) The…

Signal Processing · Electrical Eng. & Systems 2024-06-19 Navid Mohammadi Foumani , Geoffrey Mackellar , Soheila Ghane , Saad Irtza , Nam Nguyen , Mahsa Salehi

We present a modular framework for articulatory animation synthesis using speech motion capture data obtained with electromagnetic articulography (EMA). Adapting a skeletal animation approach, the articulatory motion data is applied to a…

Human-Computer Interaction · Computer Science 2012-03-19 Ingmar Steiner , Slim Ouni

Embodied human communication encompasses both verbal (speech) and non-verbal information (e.g., gesture and head movements). Recent advances in machine learning have substantially improved the technologies for generating synthetic versions…

Machine Learning · Computer Science 2021-01-15 Simon Alexanderson , Éva Székely , Gustav Eje Henter , Taras Kucherenko , Jonas Beskow

Self-supervised language models are very effective at predicting high-level cortical responses during language comprehension. However, the best current models of lower-level auditory processing in the human brain rely on either…

Computation and Language · Computer Science 2022-05-31 Aditya R. Vaidya , Shailee Jain , Alexander G. Huth

Decoding visual experience from brain signals offers exciting possibilities for neuroscience and interpretable AI. While EEG is accessible and temporally precise, its limitations in spatial detail hinder image reconstruction. Our model…

Computer Vision and Pattern Recognition · Computer Science 2025-07-11 Arshak Rezvani , Ali Akbari , Kosar Sanjar Arani , Maryam Mirian , Emad Arasteh , Martin J. McKeown

When virtual agents interact with humans, gestures are crucial to delivering their intentions with speech. Previous multimodal co-speech gesture generation models required encoded features of all modalities to generate gestures. If some…

Computer Vision and Pattern Recognition · Computer Science 2023-05-26 Gwantae Kim , Seonghyeok Noh , Insung Ham , Hanseok Ko

We propose a novel Multi-Scale Spectrogram (MSS) modelling approach to synthesise speech with an improved coarse and fine-grained prosody. We present a generic multi-scale spectrogram prediction mechanism where the system first predicts…

Audio and Speech Processing · Electrical Eng. & Systems 2021-07-01 Ammar Abbas , Bajibabu Bollepalli , Alexis Moinet , Arnaud Joly , Penny Karanasou , Peter Makarov , Simon Slangens , Sri Karlapati , Thomas Drugman

The tongue's intricate 3D structure, comprising localized functional units, plays a crucial role in the production of speech. When measured using tagged MRI, these functional units exhibit cohesive displacements and derived quantities that…

Speech Self-Supervised Learning (SSL) has demonstrated considerable efficacy in various downstream tasks. Nevertheless, prevailing self-supervised models often overlook the incorporation of emotion-related prior information, thereby…

Audio and Speech Processing · Electrical Eng. & Systems 2024-06-12 Rui Liu , Zening Ma