中文
相关论文

相关论文: Masked Autoencoders Are Articulatory Learners

200 篇论文

Articulatory-to-acoustic inversion strongly depends on the type of data used. While most previous studies rely on EMA, which is limited by the number of sensors and restricted to accessible articulators, we propose an approach aiming at a…

音频与语音处理 · 电气工程与系统科学 2026-03-31 Sofiane Azzouz , Pierre-André Vuissoz , Yves Laprie

This paper proposes a unified deep speaker embedding framework for modeling speech data with different sampling rates. Considering the narrowband spectrogram as a sub-image of the wideband spectrogram, we tackle the joint modeling problem…

音频与语音处理 · 电气工程与系统科学 2020-12-02 Weicheng Cai , Ming Li

ASR models often suffer from a long-form deletion problem where the model predicts sequential blanks instead of words when transcribing a lengthy audio (in the order of minutes or hours). From the perspective of a user or downstream system…

The amount of articulatory data available for training deep learning models is much less compared to acoustic speech data. In order to improve articulatory-to-acoustic synthesis performance in these low-resource settings, we propose a…

音频与语音处理 · 电气工程与系统科学 2024-12-19 Peter Wu , Bohan Yu , Kevin Scheck , Alan W Black , Aditi S. Krishnapriyan , Irene Y. Chen , Tanja Schultz , Shinji Watanabe , Gopala K. Anumanchipalli

Generative deep neural networks are widely used for speech synthesis, but most existing models directly generate waveforms or spectral outputs. Humans, however, produce speech by controlling articulators, which results in the production of…

声音 · 计算机科学 2023-05-10 Gašper Beguš , Alan Zhou , Peter Wu , Gopala K Anumanchipalli

In the articulatory synthesis task, speech is synthesized from input features containing information about the physical behavior of the human vocal tract. This task provides a promising direction for speech synthesis research, as the…

音频与语音处理 · 电气工程与系统科学 2022-09-15 Peter Wu , Shinji Watanabe , Louis Goldstein , Alan W Black , Gopala K. Anumanchipalli

Recent self-supervised learning (SSL) models have proven to learn rich representations of speech, which can readily be utilized by diverse downstream tasks. To understand such utilities, various analyses have been done for speech SSL models…

音频与语音处理 · 电气工程与系统科学 2023-07-24 Cheol Jun Cho , Peter Wu , Abdelrahman Mohamed , Gopala K. Anumanchipalli

Although many previous studies have carried out multimodal learning with real-time MRI data that captures the audio-visual kinematics of the vocal tract during speech, these studies have been limited by their reliance on multi-speaker…

Distant speech recognition is a challenge, particularly due to the corruption of speech signals by reverberation caused by large distances between the speaker and microphone. In order to cope with a wide range of reverberations in…

计算与语言 · 计算机科学 2016-08-18 Jeehye Lee , Myungin Lee , Joon-Hyuk Chang

Although mask-based beamforming is a powerful speech enhancement approach, it often requires manual parameter tuning to handle moving speakers. Recently, this approach was augmented with an attention-based spatial covariance matrix…

音频与语音处理 · 电气工程与系统科学 2024-06-18 Marvin Tammen , Tsubasa Ochiai , Marc Delcroix , Tomohiro Nakatani , Shoko Araki , Simon Doclo

Human infants face a formidable challenge in speech acquisition: mapping extremely variable acoustic inputs into appropriate articulatory movements without explicit instruction. We present a computational model that addresses the…

音频与语音处理 · 电气工程与系统科学 2025-09-16 Marvin Lavechin , Thomas Hueber

Speech sounds of spoken language are obtained by varying configuration of the articulators surrounding the vocal tract. They contain abundant information that can be utilized to better understand the underlying mechanism of human speech…

图像与视频处理 · 电气工程与系统科学 2021-06-17 Laxmi Pandey , Ahmed Sabbir Arif

In this paper, we combine Hidden Markov Models (HMMs) with i-vector extractors to address the problem of text-dependent speaker recognition with random digit strings. We employ digit-specific HMMs to segment the utterances into digits, to…

音频与语音处理 · 电气工程与系统科学 2019-07-16 Nooshin Maghsoodi , Hossein Sameti , Hossein Zeinali , Themos~Stafylakis

Articulatory acoustic inversion aims to reconstruct the complete geometry of the vocal tract from the speech signal. In this paper, we present a comparative study of several levels of phonetic segmentation accuracy, together with a…

音频与语音处理 · 电气工程与系统科学 2026-03-13 Sofiane Azzouz , Pierre-André Vuissoz , Yves Laprie

The articulatory geometric configurations of the vocal tract and the acoustic properties of the resultant speech sound are considered to have a strong causal relationship. This paper aims at finding a joint latent representation between the…

音频与语音处理 · 电气工程与系统科学 2020-10-02 Pramit Saha , Sidney Fels

Deep neural network based speaker embeddings, such as x-vectors, have been shown to perform well in text-independent speaker recognition/verification tasks. In this paper, we use simple classifiers to investigate the contents encoded by…

音频与语音处理 · 电气工程与系统科学 2020-06-16 Desh Raj , David Snyder , Daniel Povey , Sanjeev Khudanpur

Speech sounds are produced as the coordinated movement of the speaking organs. There are several available methods to model the relation of articulatory movements and the resulting speech signal. The reverse problem is often called as…

声音 · 计算机科学 2019-04-16 Dagoberto Porras , Alexander Sepúlveda-Sepúlveda , Tamás Gábor Csapó

Although commercial and open-source software exist to reconstruct a static object from a sequence recorded with an RGB-D sensor, there is a lack of tools that build rigged models of articulated objects that deform realistically and can be…

计算机视觉与模式识别 · 计算机科学 2016-09-12 Dimitrios Tzionas , Juergen Gall

We present a novel end-to-end identity-agnostic face reenactment system, MaskRenderer, that can generate realistic, high fidelity frames in real-time. Although recent face reenactment works have shown promising results, there are still…

计算机视觉与模式识别 · 计算机科学 2023-09-12 Tina Behrouzi , Atefeh Shahroudnejad , Payam Mousavi

This paper studies a simple extension of image-based Masked Autoencoders (MAE) to self-supervised representation learning from audio spectrograms. Following the Transformer encoder-decoder design in MAE, our Audio-MAE first encodes audio…