中文
相关论文

相关论文: emg2speech: Synthesizing speech from electromyogra…

200 篇论文

Structure-informed protein representation learning is essential for effective protein function annotation and \textit{de novo} design. However, the presence of inherent noise in both crystal and AlphaFold-predicted structures poses…

生物大分子 · 定量生物学 2025-03-25 Zhongyue Zhang , Runze Ma , Yanjie Huang , Shuangjia Zheng

In this paper we investigate continuous speech recognition using electroencephalography (EEG) features using recently introduced end-to-end transformer based automatic speech recognition (ASR) model. Our results demonstrate that transformer…

音频与语音处理 · 电气工程与系统科学 2020-05-06 Gautam Krishna , Co Tran , Mason Carnahan , Ahmed H Tewfik

Although many previous studies have carried out multimodal learning with real-time MRI data that captures the audio-visual kinematics of the vocal tract during speech, these studies have been limited by their reliance on multi-speaker…

We explore surface electromyography (sEMG) as a non-invasive input modality for mapping muscle activity to keyboard inputs, targeting immersive typing in next-generation human-computer interaction (HCI). This is especially relevant for…

人机交互 · 计算机科学 2025-11-25 Kunwoo Lee , Dhivya Sreedhar , Pushkar Saraf , Chaeeun Lee , Kateryna Shapovalenko

Articulatory-to-acoustic (A2A) synthesis refers to the generation of audible speech from captured movement of the speech articulators. This technique has numerous applications, such as restoring oral communication to people who cannot…

Decoding language from brain dynamics is an important open direction in the realm of brain-computer interface (BCI), especially considering the rapid growth of large language models. Compared to invasive-based signals which require…

计算与语言 · 计算机科学 2024-06-04 Yiqian Yang , Yiqun Duan , Qiang Zhang , Hyejeong Jo , Jinni Zhou , Won Hee Lee , Renjing Xu , Hui Xiong

A lot of work has been done to build text-based language models for performing different NLP tasks, but not much research has been done in the case of audio-based language models. This paper proposes a Convolutional Autoencoder based neural…

计算与语言 · 计算机科学 2020-09-30 Prakamya Mishra , Pranav Mathur

Speech comprehension is an involuntary task for the healthy human brain, yet the understanding of the mechanisms underlying this brain functionality remains obscure. In this paper, we aim to quantify the role of acoustic and semantic…

音频与语音处理 · 电气工程与系统科学 2025-07-31 Sai Samrat Kankanala , Akshara Soman , Sriram Ganapathy

End-to-end speech summarization (E2E SSum) directly summarizes input speech into easy-to-read short sentences with a single model. This approach is promising because it, in contrast to the conventional cascade approach, can utilize full…

Spoken language understanding, which extracts intents and/or semantic concepts in utterances, is conventionally formulated as a post-processing of automatic speech recognition. It is usually trained with oracle transcripts, but needs to…

声音 · 计算机科学 2020-07-30 Viet-Trung Dang , Tianyu Zhao , Sei Ueno , Hirofumi Inaguma , Tatsuya Kawahara

Recent advances in machine learning and the availability of articulatory datasets allow vocal tract synthesis to be conditioned on phonetic sequences, a primary task of articulatory speech synthesis. However, quality assessment needs a…

计算与语言 · 计算机科学 2026-05-21 Vinicius Ribeiro , Yves Laprie

Surface electromyography (sEMG) signals exhibit substantial inter-subject variability and are highly susceptible to noise, posing challenges for robust and interpretable decoding. To address these limitations, we propose a discrete…

信号处理 · 电气工程与系统科学 2026-03-02 Yuepeng Chen , Kaili Zheng , Ji Wu , Zhuangzhuang Li , Ye Ma , Dongwei Liu , Chenyi Guo , Xiangling Fu

Much recent work on Spoken Language Understanding (SLU) falls short in at least one of three ways: models were trained on oracle text input and neglected the Automatics Speech Recognition (ASR) outputs, models were trained to predict only…

计算与语言 · 计算机科学 2020-11-13 Cheng-I Lai , Jin Cao , Sravan Bodapati , Shang-Wen Li

In this work, we focus on source tracing of synthetic speech generation systems (STSGS). Each source embeds distinctive paralinguistic features--such as pitch, tone, rhythm, and intonation--into their synthesized speech, reflecting the…

Despite being trained exclusively on speech data, speech foundation models (SFMs) like Whisper have shown impressive performance in non-speech tasks such as audio classification. This is partly because speech shares some common traits with…

音频与语音处理 · 电气工程与系统科学 2024-10-17 Orchid Chetia Phukan , Swarup Ranjan Behera , Girish , Mohd Mujtaba Akhtar , Arun Balaji Buduru , Rajesh Sharma

In this paper, we present EH-MAM (Easy-to-Hard adaptive Masked Acoustic Modeling), a novel self-supervised learning approach for speech representation learning. In contrast to the prior methods that use random masking schemes for Masked…

声音 · 计算机科学 2024-10-18 Ashish Seth , Ramaneswaran Selvakumar , S Sakshi , Sonal Kumar , Sreyan Ghosh , Dinesh Manocha

Speech is a means of communication which relies on both audio and visual information. The absence of one modality can often lead to confusion or misinterpretation of information. In this paper we present an end-to-end temporal model capable…

音频与语音处理 · 电气工程与系统科学 2019-06-17 Konstantinos Vougioukas , Pingchuan Ma , Stavros Petridis , Maja Pantic

We propose a novel self-supervised embedding to learn how actions sound from narrated in-the-wild egocentric videos. Whereas existing methods rely on curated data with known audio-visual correspondence, our multimodal contrastive-consensus…

计算机视觉与模式识别 · 计算机科学 2024-04-09 Changan Chen , Kumar Ashutosh , Rohit Girdhar , David Harwath , Kristen Grauman

Speech synthesis from intracranial EEG (iEEG) signals offers a promising avenue for restoring communication in individuals with severe speech impairments. However, achieving intelligible and natural speech remains challenging due to…

声音 · 计算机科学 2025-08-06 Mohammed Salah Al-Radhi , Géza Németh , Branislav Gerazov

Some recent studies have demonstrated the feasibility of single-stage neural text-to-speech, which does not need to generate mel-spectrograms but generates the raw waveforms directly from the text. Single-stage text-to-speech often faces…

声音 · 计算机科学 2022-07-14 Zhengxi Liu , Qiao Tian , Chenxu Hu , Xudong Liu , Menglin Wu , Yuping Wang , Hang Zhao , Yuxuan Wang