English
Related papers

Related papers: Speech animation using electromagnetic articulogra…

200 papers

Articulatory-to-acoustic (forward) mapping is a technique to predict speech using various articulatory acquisition techniques (e.g. ultrasound tongue imaging, lip video). Real-time MRI (rtMRI) of the vocal tract has not been used before for…

Audio and Speech Processing · Electrical Eng. & Systems 2020-08-04 Tamás Gábor Csapó

Human motion diffusion models can synthesize action sequences from text, but controlling motion intensity remains challenging. Existing approaches rely on effort-related adverbs, which are ambiguous and fail to capture quantitative aspects…

Computer Vision and Pattern Recognition · Computer Science 2026-05-26 Joshua Siy , Huakun Liu , Yutaro Hirao , Monica Perusquia-Hernandez , Hideaki Uchiyama , Kiyoshi Kiyokawa

A 3D biomechanical dynamical model of human tongue is presented, that is elaborated in the aim to test hypotheses about speech motor control. Tissue elastic properties are accounted for in Finite Element Modeling (FEM). The FEM mesh was…

Medical Physics · Physics 2007-05-23 Jean-Michel Gerard , Reiner Wilhelms-Tricarico , Pascal Perrier , Yohan Payan

We present a multilinear statistical model of the human tongue that captures anatomical and tongue pose related shape variations separately. The model is derived from 3D magnetic resonance imaging data of 11 speakers sustaining speech…

Computer Vision and Pattern Recognition · Computer Science 2018-04-18 Alexander Hewer , Stefanie Wuhrer , Ingmar Steiner , Korin Richmond

This paper presents a novel approach for generating 3D talking heads from raw audio inputs. Our method grounds on the idea that speech related movements can be comprehensively and efficiently described by the motion of a few control points…

Computer Vision and Pattern Recognition · Computer Science 2023-07-27 Federico Nocentini , Claudio Ferrari , Stefano Berretti

This paper describes an extension of the two-dimensional dynamic articulatory model DYNARTmo by integrating an internal three-dimensional representation of the palatal dome to estimate tongue-palate contact areas from midsagittal tongue…

Computation and Language · Computer Science 2025-08-12 Bernd J. Kröger

We present DYNARTmo, a dynamic articulatory model designed to visualize speech articulation processes in a two-dimensional midsagittal plane. The model builds upon the UK-DYNAMO framework and integrates principles of articulatory…

Computation and Language · Computer Science 2025-11-07 Bernd J. Kröger

This paper describes a technique for using magnetic motion capture data to determine the joint parameters of an articulated hierarchy. This technique makes it possible to determine limb lengths, joint locations, and sensor placement for a…

Graphics · Computer Science 2023-03-21 James F. O'Brien , Robert E. Bodenheimer , Gabriel J. Brostow , Jessica K. Hodgins

We present a neuromuscular speech interface that translates silently voiced articulations directly into text. We record surface electromyographic (EMG) signals from multiple articulatory sites on the face and neck as participants silently…

Audio and Speech Processing · Electrical Eng. & Systems 2026-04-21 Harshavardhana T. Gowda , Lee M. Miller

Speech-driven facial animation requires accurate correspondence between acoustic signals and facial motion, especially for articulation-related mouth movements. However, directly mapping speech audio to facial coefficients often overlooks…

Computer Vision and Pattern Recognition · Computer Science 2026-05-11 Kai Zheng , Zejian Kang , Rui Mao , Hongyuan Zou , Yuanchen Fei , Xuanyang Xu , Xiangru Huang

Previous initial research has already been carried out to propose speech-based BCI using brain signals (e.g. non-invasive EEG and invasive sEEG / ECoG), but there is a lack of combined methods that investigate non-invasive brain,…

Medical Physics · Physics 2023-10-19 Tamás Gábor Csapó , Frigyes Viktor Arthur , Péter Nagy , Ádám Boncz

In this paper, we present our first experiments in text-to-articulation prediction, using ultrasound tongue image targets. We extend a traditional (vocoder-based) DNN-TTS framework with predicting PCA-compressed ultrasound images, of which…

Audio and Speech Processing · Electrical Eng. & Systems 2021-07-13 Tamás Gábor Csapó

In this paper, we consider the task of digitally voicing silent speech, where silently mouthed words are converted to audible speech based on electromyography (EMG) sensor measurements that capture muscle impulses. While prior work has…

Audio and Speech Processing · Electrical Eng. & Systems 2020-10-08 David Gaddy , Dan Klein

Speech emotion recognition is a challenging problem because human convey emotions in subtle and complex ways. For emotion recognition on human speech, one can either extract emotion related features from audio signals or employ speech…

Computation and Language · Computer Science 2020-04-06 Haiyang Xu , Hui Zhang , Kun Han , Yun Wang , Yiping Peng , Xiangang Li

Motions of virtual characters in movies or video games are typically generated by recording actors using motion capturing methods. Animations generated this way often need postprocessing, such as improving the periodicity of cyclic…

Computer Vision and Pattern Recognition · Computer Science 2015-02-27 Martin Bauer , Markus Eslitzbichler , Markus Grasmair

Speech generation and enhancement based on articulatory movements facilitate communication when the scope of verbal communication is absent, e.g., in patients who have lost the ability to speak. Although various techniques have been…

Sound · Computer Science 2023-11-29 Li-Chin Chen , Po-Hsun Chen , Richard Tzong-Han Tsai , Yu Tsao

Cross-corpus speech emotion recognition (SER) plays a vital role in numerous practical applications. Traditional approaches to cross-corpus emotion transfer often concentrate on adapting acoustic features to align with different corpora,…

Sound · Computer Science 2024-12-31 Shreya G. Upadhyay , Ali N. Salman , Carlos Busso , Chi-Chun Lee

This work focuses on full-body co-speech gesture generation. Existing methods typically employ an autoregressive model accompanied by vector-quantized tokens for gesture generation, which results in information loss and compromises the…

Graphics · Computer Science 2025-03-19 Binjie Liu , Lina Liu , Sanyi Zhang , Songen Gu , Yihao Zhi , Tianyi Zhu , Lei Yang , Long Ye

We estimate articulatory movements in speech production from different modalities - acoustics and phonemes. Acoustic-to articulatory inversion (AAI) is a sequence-to-sequence task. On the other hand, phoneme to articulatory (PTA) motion…

Audio and Speech Processing · Electrical Eng. & Systems 2021-06-15 Sathvik Udupa , Anwesha Roy , Abhayjeet Singh , Aravind Illa , Prasanta Kumar Ghosh

Quantitative measurement of functional and anatomical traits of 4D tongue motion in the course of speech or other lingual behaviors remains a major challenge in scientific research and clinical applications. Here, we introduce a statistical…

Computer Vision and Pattern Recognition · Computer Science 2018-09-18 Jonghye Woo , Fangxu Xing , Maureen Stone , Jordan Green , Timothy G. Reese , Thomas J. Brady , Van J. Wedeen , Jerry L. Prince , Georges El Fakhri