中文
相关论文

相关论文: Speech animation using electromagnetic articulogra…

200 篇论文

Codec Avatars are a recent class of learned, photorealistic face models that accurately represent the geometry and texture of a person in 3D (i.e., for virtual reality), and are almost indistinguishable from video. In this paper we describe…

计算机视觉与模式识别 · 计算机科学 2020-08-13 Alexander Richard , Colin Lea , Shugao Ma , Juergen Gall , Fernando de la Torre , Yaser Sheikh

The electroencephalography (EEG) signals recorded in parallel with speech are used to perform isolated and continuous speech recognition. During speaking process, one also hears his or her own speech and this speech perception is also…

音频与语音处理 · 电气工程与系统科学 2020-06-03 Gautam Krishna , Co Tran , Mason Carnahan , Ahmed Tewfik

Although significant progress has been made in the field of speech-driven 3D facial animation recently, the speech-driven animation of an indispensable facial component, eye gaze, has been overlooked by recent research. This is primarily…

计算机视觉与模式识别 · 计算机科学 2026-01-06 Yixiang Zhuang , Chunshan Ma , Yao Cheng , Xuan Cheng , Jing Liao , Juncong Lin

We propose a novel robust and efficient Speech-to-Animation (S2A) approach for synchronized facial animation generation in human-computer interaction. Compared with conventional approaches, the proposed approach utilizes phonetic…

多媒体 · 计算机科学 2022-04-07 Liyang Chen , Zhiyong Wu , Jun Ling , Runnan Li , Xu Tan , Sheng Zhao

Unlike phoneme sequences, movements of speech articulators (lips, tongue, jaw, velum) and the resultant acoustic signal are known to encode not only the linguistic message but also carry para-linguistic information. While several works…

音频与语音处理 · 电气工程与系统科学 2020-02-20 Abhayjeet Singh , Aravind Illa , Prasanta Kumar Ghosh

Lip-reading is to utilize the visual information of the speaker's lip movements to recognize words and sentences. Existing event-based lip-reading solutions integrate different frame rate branches to learn spatio-temporal features of…

计算机视觉与模式识别 · 计算机科学 2025-04-14 Wenhao Zhang , Jun Wang , Yong Luo , Lei Yu , Wei Yu , Zheng He , Jialie Shen

This paper introduces a novel combination of two tasks, previously treated separately: acoustic-to-articulatory speech inversion (AAI) and phoneme-to-articulatory (PTA) motion estimation. We refer to this joint task as acoustic…

This paper presents a generic method for generating full facial 3D animation from speech. Existing approaches to audio-driven facial animation exhibit uncanny or static upper face animation, fail to produce accurate and plausible…

计算机视觉与模式识别 · 计算机科学 2022-05-23 Alexander Richard , Michael Zollhoefer , Yandong Wen , Fernando de la Torre , Yaser Sheikh

Acoustic emotion recognition aims to categorize the affective state of the speaker and is still a difficult task for machine learning models. The difficulties come from the scarcity of training data, general subjectivity in emotion…

计算与语言 · 计算机科学 2018-04-02 Egor Lakomkin , Cornelius Weber , Sven Magg , Stefan Wermter

We present an end-to-end text-to-speech (TTS) synthesis system that generates audio and synchronized tongue motion directly from text. This is achieved by adapting a 3D model of the tongue surface to an articulatory dataset and training a…

人机交互 · 计算机科学 2018-04-17 Ingmar Steiner , Sébastien Le Maguer , Alexander Hewer

Human speech is often accompanied by body gestures including arm and hand gestures. We present a method that reenacts a high-quality video with gestures matching a target speech audio. The key idea of our method is to split and re-assemble…

计算机视觉与模式识别 · 计算机科学 2022-07-26 Yang Zhou , Jimei Yang , Dingzeyu Li , Jun Saito , Deepali Aneja , Evangelos Kalogerakis

Mobile health (mHealth) systems help researchers monitor and care for patients in real-world settings. Studies utilizing mHealth applications use Ecological Momentary Assessment (EMAs), passive sensing, and contextual features to develop…

Many speech segments in movies are re-recorded in a studio during postproduction, to compensate for poor sound quality as recorded on location. Manual alignment of the newly-recorded speech with the original lip movements is a tedious task.…

计算机视觉与模式识别 · 计算机科学 2018-08-21 Tavi Halperin , Ariel Ephrat , Shmuel Peleg

Most state-of-the-art sign language models are trained on interpreter or isolated vocabulary data, which overlooks the variability that characterizes natural dialogue. However, human communication dynamically adapts to contexts and…

计算与语言 · 计算机科学 2025-10-29 Saki Imai , Lee Kezar , Laurel Aichler , Mert Inan , Erin Walker , Alicia Wooten , Lorna Quandt , Malihe Alikhani

In this work, we tackle the challenge of enhancing the realism and expressiveness in talking head video generation by focusing on the dynamic and nuanced relationship between audio cues and facial movements. We identify the limitations of…

计算机视觉与模式识别 · 计算机科学 2024-08-09 Linrui Tian , Qi Wang , Bang Zhang , Liefeng Bo

Surface Electromyography (sEMG) provides vital insights into muscle function, but it can be noisy and challenging to acquire. Inertial Measurement Units (IMUs) provide a robust and wearable alternative to motion capture systems. This paper…

机器学习 · 计算机科学 2025-11-24 Shubhranil Basak , Mada Hemanth , Madhav Rao

Efficiently digitizing high-fidelity animatable human avatars from videos is a challenging and active research topic. Recent volume rendering-based neural representations open a new way for human digitization with their friendly usability…

计算机视觉与模式识别 · 计算机科学 2023-03-24 Xiaoke Huang , Yiji Cheng , Yansong Tang , Xiu Li , Jie Zhou , Jiwen Lu

Ecological momentary assessment (EMA) data have a broad base of application in the study of time trends and relations. In EMA studies, there are a number of design considerations which influence the analysis of the data. One general…

Recently audio-driven talking face video generation has attracted considerable attention. However, very few researches address the issue of emotional editing of these talking face videos with continuously controllable expressions, which is…

计算机视觉与模式识别 · 计算机科学 2023-11-29 Zhiyao Sun , Yu-Hui Wen , Tian Lv , Yanan Sun , Ziyang Zhang , Yaoyuan Wang , Yong-Jin Liu

The task of emotion recognition in conversations (ERC) benefits from the availability of multiple modalities, as provided, for example, in the video-based Multimodal EmotionLines Dataset (MELD). However, only a few research approaches use…

音频与语音处理 · 电气工程与系统科学 2023-08-16 Hugo Carneiro , Cornelius Weber , Stefan Wermter