中文
相关论文

相关论文: Discovering phoneme-specific critical articulators…

200 篇论文

As a first step towards a complete computational model of speech learning involving perception-production loops, we investigate the forward mapping between pseudo-motor commands and articulatory trajectories. Two phonological feature sets,…

音频与语音处理 · 电气工程与系统科学 2024-08-09 Angelo Ortiz Tandazo , Thomas Schatz , Thomas Hueber , Emmanuel Dupoux

Automatic phonemic transcription tools are useful for low-resource language documentation. However, due to the lack of training sets, only a tiny fraction of languages have phonemic transcription tools. Fortunately, multilingual acoustic…

计算与语言 · 计算机科学 2020-02-28 Xinjian Li , Siddharth Dalmia , David R. Mortensen , Juncheng Li , Alan W Black , Florian Metze

The way infants use auditory cues to learn to speak despite the acoustic mismatch of their vocal apparatus is a hot topic of scientific debate. The simulation of early vocal learning using articulatory speech synthesis offers a way towards…

音频与语音处理 · 电气工程与系统科学 2021-04-05 Branislav Gerazov , Daniel van Niekerk , Anqi Xu , Paul Konstantin Krug , Peter Birkholz , Yi Xu

We introduce a self-supervised speech pre-training method called TERA, which stands for Transformer Encoder Representations from Alteration. Recent approaches often learn by using a single auxiliary task like contrastive prediction,…

音频与语音处理 · 电气工程与系统科学 2021-08-05 Andy T. Liu , Shang-Wen Li , Hung-yi Lee

We present EMPHASIS, an emotional phoneme-based acoustic model for speech synthesis system. EMPHASIS includes a phoneme duration prediction model and an acoustic parameter prediction model. It uses a CBHG-based regression network to model…

音频与语音处理 · 电气工程与系统科学 2018-06-27 Hao Li , Yongguo Kang , Zhenyu Wang

Electromagnetic articulography (EMA) captures the position and orientation of a number of markers, attached to the articulators, during speech. As such, it performs the same function for speech that conventional motion capture does for…

人机交互 · 计算机科学 2013-11-01 Ingmar Steiner , Korin Richmond , Slim Ouni

Direct acoustics-to-word (A2W) systems for end-to-end automatic speech recognition are simpler to train, and more efficient to decode with, than sub-word systems. However, A2W systems can have difficulties at training time when data is…

计算与语言 · 计算机科学 2019-04-01 Shane Settle , Kartik Audhkhasi , Karen Livescu , Michael Picheny

In this paper we introduce attention-regression model to demonstrate predicting acoustic features from electroencephalography (EEG) features recorded in parallel with spoken sentences. First we demonstrate predicting acoustic features…

音频与语音处理 · 电气工程与系统科学 2020-05-05 Gautam Krishna , Co Tran , Mason Carnahan , Ahmed Tewfik

It is increasingly considered that human speech perception and production both rely on articulatory representations. In this paper, we investigate whether this type of representation could improve the performances of a deep generative model…

声音 · 计算机科学 2021-04-08 Marc-Antoine Georges , Laurent Girin , Jean-Luc Schwartz , Thomas Hueber

Generative deep neural networks are widely used for speech synthesis, but most existing models directly generate waveforms or spectral outputs. Humans, however, produce speech by controlling articulators, which results in the production of…

声音 · 计算机科学 2023-05-10 Gašper Beguš , Alan Zhou , Peter Wu , Gopala K Anumanchipalli

Prior works have investigated the use of articulatory features as complementary representations for automatic speech recognition (ASR), but their use was largely confined to shallow acoustic models. In this work, we revisit articulatory…

音频与语音处理 · 电气工程与系统科学 2025-10-13 Ahmed Adel Attia , Jing Liu , Carol Espy Wilson

Real-time Magnetic Resonance Imaging (rtMRI) visualizes vocal tract action, offering a comprehensive window into speech articulation. However, its signals are high dimensional and noisy, hindering interpretation. We investigate compact…

图像与视频处理 · 电气工程与系统科学 2026-01-30 Jay Park , Hong Nguyen , Sean Foley , Jihwan Lee , Yoonjeong Lee , Dani Byrd , Shrikanth Narayanan

Models of acoustic word embeddings (AWEs) learn to map variable-length spoken word segments onto fixed-dimensionality vector representations such that different acoustic exemplars of the same word are projected nearby in the embedding…

计算与语言 · 计算机科学 2022-09-20 Badr M. Abdullah , Bernd Möbius , Dietrich Klakow

We present a novel neural encoder system for acoustic-to-articulatory inversion. We leverage the Pink Trombone voice synthesizer that reveals articulatory parameters (e.g tongue position and vocal cord configuration). Our system is designed…

音频与语音处理 · 电气工程与系统科学 2024-06-21 Mateo Cámara , Fernando Marcos , José Luis Blanco

We address the problem of reconstructing articulatory movements, given audio and/or phonetic labels. The scarce availability of multi-speaker articulatory data makes it difficult to learn a reconstruction that generalizes to new speakers…

计算与语言 · 计算机科学 2023-09-13 Rosanna Turrisi , Raffaele Tavarone , Leonardo Badino

Speech production involves the movement of various articulators, including tongue, jaw, and lips. Estimating the movement of the articulators from the acoustics of speech is known as acoustic-to-articulatory inversion (AAI). Recently, it…

音频与语音处理 · 电气工程与系统科学 2020-06-23 Aravind Illa , Prasanta Kumar Ghosh

While speaking at different rates, articulators (like tongue, lips) tend to move differently and the enunciations are also of different durations. In the past, affine transformation and DNN have been used to transform articulatory movements…

音频与语音处理 · 电气工程与系统科学 2020-08-21 Abhayjeet Singh , Aravind Illa , Prasanta Kumar Ghosh

Acoustic-to-articulatory inversion (AAI) involves mapping from the acoustic to the articulatory space. Signal-processing features like the MFCCs, have been widely used for the AAI task. For subjects with dysarthric speech, AAI is…

音频与语音处理 · 电气工程与系统科学 2024-02-13 Sarthak Kumar Maharana , Krishna Kamal Adidam , Shoumik Nandi , Ajitesh Srivastava

We propose an attention-enabled encoder-decoder model for the problem of grapheme-to-phoneme conversion. Most previous work has tackled the problem via joint sequence models that require explicit alignments for training. In contrast, the…

计算与语言 · 计算机科学 2016-10-21 Shubham Toshniwal , Karen Livescu

Phonemic segmentation of speech is a critical step of speech recognition systems. We propose a novel unsupervised algorithm based on sequence prediction models such as Markov chains and recurrent neural network. Our approach consists in…

计算与语言 · 计算机科学 2017-05-30 Paul Michel , Okko Räsänen , Roland Thiollière , Emmanuel Dupoux