English
Related papers

Related papers: Speech animation using electromagnetic articulogra…

200 papers

As 3D facial avatars become more widely used for communication, it is critical that they faithfully convey emotion. Unfortunately, the best recent methods that regress parametric 3D face models from monocular images are unable to capture…

Computer Vision and Pattern Recognition · Computer Science 2022-04-26 Radek Danecek , Michael J. Black , Timo Bolkart

Speech-driven 3D facial animation has recently garnered attention due to its cost-effective usability in multimedia production. However, most current advances overlook the intelligibility of lip movements, limiting the realism of facial…

Computer Vision and Pattern Recognition · Computer Science 2024-07-02 Han EunGi , Oh Hyun-Bin , Kim Sung-Bin , Corentin Nivelet Etcheberry , Suekyeong Nam , Janghoon Joo , Tae-Hyun Oh

This article describes the development of a platform designed to visualize the 3D motion of the tongue using ultrasound image sequences. An overview of the system design is given and promising results are presented. Compared to the analysis…

Computer Vision and Pattern Recognition · Computer Science 2016-05-20 Kele Xu , Yin Yang , Aurore Jaumard-Hakoun , Clemence Leboullenger , Gerard Dreyfus , Pierre Roussel , Maureen Stone , Bruce Denby

Audio-driven 3D facial animation has been widely explored, but achieving realistic, human-like performance is still unsolved. This is due to the lack of available 3D datasets, models, and standard evaluation metrics. To address this, we…

Computer Vision and Pattern Recognition · Computer Science 2019-05-09 Daniel Cudeiro , Timo Bolkart , Cassidy Laidlaw , Anurag Ranjan , Michael J. Black

Speech-driven 3D facial animation has been widely studied, yet there is still a gap to achieving realism and vividness due to the highly ill-posed nature and scarcity of audio-visual data. Existing works typically formulate the cross-modal…

Computer Vision and Pattern Recognition · Computer Science 2023-04-04 Jinbo Xing , Menghan Xia , Yuechen Zhang , Xiaodong Cun , Jue Wang , Tien-Tsin Wong

Several approaches exist for the recording of articulatory movements, such as eletromagnetic and permanent magnetic articulagraphy, ultrasound tongue imaging and surface electromyography. Although magnetic resonance imaging (MRI) is more…

Sound · Computer Science 2021-04-26 Yide Yu , Amin Honarmandi Shandiz , László Tóth

Multimodal learning has been proven to be an effective method to improve speech enhancement (SE) performance, especially in challenging situations such as low signal-to-noise ratios, speech noise, or unseen noise types. In previous studies,…

Audio and Speech Processing · Electrical Eng. & Systems 2022-02-15 Kuan-Chen Wang , Kai-Chun Liu , Hsin-Min Wang , Yu Tsao

We study diffusion-based speech enhancement using a Schrodinger bridge formulation and extend the EDM2 framework to this setting. We employ time-dependent preconditioning of network inputs and outputs to stabilize training and explore two…

Audio and Speech Processing · Electrical Eng. & Systems 2026-01-30 Julius Richter , Danilo de Oliveira , Timo Gerkmann

Real-time Magnetic Resonance Imaging (rtMRI) visualizes vocal tract action, offering a comprehensive window into speech articulation. However, its signals are high dimensional and noisy, hindering interpretation. We investigate compact…

Image and Video Processing · Electrical Eng. & Systems 2026-01-30 Jay Park , Hong Nguyen , Sean Foley , Jihwan Lee , Yoonjeong Lee , Dani Byrd , Shrikanth Narayanan

Ecological momentary assessment (EMA) is used to evaluate subjects' behaviors and moods in their natural environments, yet collecting real-time and self-report data with EMA is challenging due to user burden. Integrating voice into EMA data…

Human-Computer Interaction · Computer Science 2024-07-09 Chen Chen , Khalil Mrini , Kemeberly Charles , Ella T. Lifset , Michael Hogarth , Alison A. Moore , Nadir Weibel , Emilia Farcas

We introduce EMMA, a physics-informed multimodal framework that recovers all identifiable dynamical parameters of a system directly from raw video, audio, and image-based time-series observations. Unlike prior video-only approaches that…

Computer Vision and Pattern Recognition · Computer Science 2026-05-26 Farhat Shaikh , Ayan Banerjee , Sandeep Gupta

This paper proposes a novel 3D speech-to-animation (STA) generation framework designed to address the shortcomings of existing models in producing diverse and emotionally resonant animations. Current STA models often generate animations…

Computer Vision and Pattern Recognition · Computer Science 2024-11-27 Xulong Zhang , Xiaoyang Qu , Haoxiang Shi , Chunguang Xiao , Jianzong Wang

This paper presents Articulatory-WaveNet, a new approach for acoustic-to-articulator inversion. The proposed system uses the WaveNet speech synthesis architecture, with dilated causal convolutional layers using previous values of the…

Audio and Speech Processing · Electrical Eng. & Systems 2020-06-24 Narjes Bozorg , Michael T. Johnson

Patient movement in emission tomography deteriorates reconstruction quality because of motion blur. Gating the data improves the situation somewhat: each gate contains a movement phase which is approximately stationary. A standard method is…

Image and Video Processing · Electrical Eng. & Systems 2020-02-24 Ozan Öktem , Camille Pouchol , Olivier Verdier

We introduce the Efficient Monotonic Multihead Attention (EMMA), a state-of-the-art simultaneous translation model with numerically-stable and unbiased monotonic alignment estimation. In addition, we present improved training and inference…

Computation and Language · Computer Science 2023-12-08 Xutai Ma , Anna Sun , Siqi Ouyang , Hirofumi Inaguma , Paden Tomasello

This paper describes the current implementation of the dynamic articulatory model DYNARTmo, which generates continuous articulator movements based on the concept of speech gestures and a corresponding gesture score. The model provides a…

Computation and Language · Computer Science 2025-11-12 Bernd J. Kröger

There are a variety of features of the human voice that can be classified as pitch, timbre, loudness, and vocal tone. It is observed in numerous incidents that human expresses their feelings using different vocal qualities when they are…

A 3D biomechanical dynamical model of human tongue is presented, which is elaborated to test in the future hypotheses about speech motor control. Tissue elastic properties are accounted for in Finite Element Modelling (FEM). The FEM mesh…

Medical Physics · Physics 2007-05-23 Jean-Michel Gerard , Pascal Perrier , Yohan Payan

Most current video MLLMs rely on uniform frame sampling and image-level encoders, resulting in inefficient data processing and limited motion awareness. To address these challenges, we introduce EMA, an Efficient Motion-Aware video MLLM…

Computer Vision and Pattern Recognition · Computer Science 2025-03-18 Zijia Zhao , Yuqi Huo , Tongtian Yue , Longteng Guo , Haoyu Lu , Bingning Wang , Weipeng Chen , Jing Liu

The relationship between muscle activity and resulting facial expressions is crucial for various fields, including psychology, medicine, and entertainment. The synchronous recording of facial mimicry and muscular activity via surface…

Computer Vision and Pattern Recognition · Computer Science 2025-03-13 Tim Büchner , Christoph Anders , Orlando Guntinas-Lichius , Joachim Denzler