English
Related papers

Related papers: ARTI-6: Towards Six-dimensional Articulatory Speec…

200 papers

Speech is produced through the coordination of vocal tract constricting organs: lips, tongue, velum, and glottis. Previous works developed Speech Inversion (SI) systems to recover acoustic-to-articulatory mappings for lip and tongue…

Audio and Speech Processing · Electrical Eng. & Systems 2025-09-01 Saba Tabatabaee , Suzanne Boyce , Liran Oren , Mark Tiede , Carol Espy-Wilson

Conversion of non-native accented speech to native (American) English has a wide range of applications such as improving intelligibility of non-native speech. Previous work on this domain has used phonetic posteriograms as the target speech…

Audio and Speech Processing · Electrical Eng. & Systems 2024-06-11 Yashish M. Siriwardena , Nathan Swedlow , Audrey Howard , Evan Gitterman , Dan Darcy , Carol Espy-Wilson , Andrea Fanelli

Current speech production systems predominantly rely on large transformer models that operate as black boxes, providing little interpretability or grounding in the physical mechanisms of human speech. We address this limitation by proposing…

Audio and Speech Processing · Electrical Eng. & Systems 2025-10-08 Akshay Anand , Chenxu Guo , Cheol Jun Cho , Jiachen Lian , Gopala Anumanchipalli

Investigating the relationship between internal tissue point motion of the tongue and oropharyngeal muscle deformation measured from tagged MRI and intelligible speech can aid in advancing speech motor control theories and developing novel…

Image and Video Processing · Electrical Eng. & Systems 2023-02-15 Xiaofeng Liu , Fangxu Xing , Jerry L. Prince , Maureen Stone , Georges El Fakhri , Jonghye Woo

Speech production involves the movement of various articulators, including tongue, jaw, and lips. Estimating the movement of the articulators from the acoustics of speech is known as acoustic-to-articulatory inversion (AAI). Recently, it…

Audio and Speech Processing · Electrical Eng. & Systems 2020-06-23 Aravind Illa , Prasanta Kumar Ghosh

Accurate analysis of speech articulation is crucial for speech analysis. However, X-Y coordinates of articulators strongly depend on the anatomy of the speakers and the variability of pellet placements, and existing methods for mapping…

Audio and Speech Processing · Electrical Eng. & Systems 2023-09-29 Ahmed Adel Attia , Mark Tiede , Carol Y. Espy-Wilson

End-to-end transformer-based automatic speech recognition (ASR) systems often capture multiple speech traits in their learned representations that are highly entangled, leading to a lack of interpretability. In this study, we propose the…

Audio and Speech Processing · Electrical Eng. & Systems 2024-11-28 Pu Wang , Hugo Van hamme

Articulatory information has been shown to be effective in improving the performance of HMM-based and DNN-based text-to-speech synthesis. Speech synthesis research focuses traditionally on text-to-speech conversion, when the input is text…

Audio and Speech Processing · Electrical Eng. & Systems 2021-07-06 Tamás Gábor Csapó , László Tóth , Gábor Gosztolya , Alexandra Markó

In this work, we incorporated acoustically derived source features, aperiodicity, periodicity and pitch as additional targets to an acoustic-to-articulatory speech inversion (SI) system. We also propose a Temporal Convolution based SI…

Audio and Speech Processing · Electrical Eng. & Systems 2022-11-01 Yashish M. Siriwardena , Carol Espy-Wilson

Recognition of speech, and in particular the ability to generalize and learn from small sets of labelled examples like humans do, depends on an appropriate representation of the acoustic input. We formulate the problem of finding robust…

Understanding the underlying relationship between tongue and oropharyngeal muscle deformation seen in tagged-MRI and intelligible speech plays an important role in advancing speech motor control theories and treatment of speech…

Accent variability has posed a huge challenge to automatic speech recognition~(ASR) modeling. Although one-hot accent vector based adaptation systems are commonly used, they require prior knowledge about the target accent and cannot handle…

Sound · Computer Science 2022-04-22 Xun Gong , Yizhou Lu , Zhikai Zhou , Yanmin Qian

Audio-visual automatic speech recognition (AV-ASR) introduces the video modality into the speech recognition process, often by relying on information conveyed by the motion of the speaker's mouth. The use of the video signal requires…

Computer Vision and Pattern Recognition · Computer Science 2021-09-21 Dmitriy Serdyuk , Otavio Braga , Olivier Siohan

Articulatory features are inherently invariant to acoustic signal distortion and have been successfully incorporated into automatic speech recognition (ASR) systems for normal speech. Their practical application to disordered speech…

Audio and Speech Processing · Electrical Eng. & Systems 2022-03-22 Shujie Hu , Shansong Liu , Xurong Xie , Mengzhe Geng , Tianzi Wang , Shoukang Hu , Mingyu Cui , Xunying Liu , Helen Meng

Dysarthric speech reconstruction (DSR) aims to transform dysarthric speech into normal speech by improving the intelligibility and naturalness. This is a challenging task especially for patients with severe dysarthria and speaking in…

Sound · Computer Science 2024-02-01 Xueyuan Chen , Yuejiao Wang , Xixin Wu , Disong Wang , Zhiyong Wu , Xunying Liu , Helen Meng

Although many previous studies have carried out multimodal learning with real-time MRI data that captures the audio-visual kinematics of the vocal tract during speech, these studies have been limited by their reliance on multi-speaker…

Machine Learning · Computer Science 2025-06-02 Sean Foley , Hong Nguyen , Jihwan Lee , Sudarsana Reddy Kadiri , Dani Byrd , Louis Goldstein , Shrikanth Narayanan

Multi-task learning (MTL) frameworks have proven to be effective in diverse speech related tasks like automatic speech recognition (ASR) and speech emotion recognition. This paper proposes a MTL framework to perform acoustic-to-articulatory…

Audio and Speech Processing · Electrical Eng. & Systems 2023-05-18 Yashish M. Siriwardena , Ganesh Sivaraman , Carol Espy-Wilson

Automatic speech recognition (ASR) is a key area in computational linguistics, focusing on developing technologies that enable computers to convert spoken language into text. This field combines linguistics and machine learning. ASR models,…

Computation and Language · Computer Science 2024-06-27 Anish Saha , A. G. Ramakrishnan

We propose a first step toward multilingual end-to-end automatic speech recognition (ASR) by integrating knowledge about speech articulators. The key idea is to leverage a rich set of fundamental units that can be defined "universally"…

Audio and Speech Processing · Electrical Eng. & Systems 2023-09-19 Hao Yen , Sabato Marco Siniscalchi , Chin-Hui Lee

Acoustic-to-Articulatory Inversion (AAI) attempts to model the inverse mapping from speech to articulation. Exact articulatory prediction from speech alone may be impossible, as speakers can choose different forms of articulation seemingly…

Sound · Computer Science 2025-06-10 Charles McGhee , Mark J. F. Gales , Kate M. Knill