中文
相关论文

相关论文: Generating Holistic 3D Human Motion from Speech

200 篇论文

Animating virtual characters with holistic co-speech gestures is a challenging but critical task. Previous systems have primarily focused on the weak correlation between audio and gestures, leading to physically unnatural outcomes that…

计算机视觉与模式识别 · 计算机科学 2025-03-18 Yongkang Cheng , Shaoli Huang

Gestures that accompany speech are an essential part of natural and efficient embodied human communication. The automatic generation of such co-speech gestures is a long-standing problem in computer animation and is considered an enabling…

图形学 · 计算机科学 2023-04-11 Simbarashe Nyatsanga , Taras Kucherenko , Chaitanya Ahuja , Gustav Eje Henter , Michael Neff

Different people have different facial expressions while speaking emotionally. A realistic facial animation system should consider such identity-specific speaking styles and facial idiosyncrasies to achieve high-degree of naturalness and…

人工智能 · 计算机科学 2023-10-27 Elif Bozkurt

Generating videos of complex human motions such as flips, cartwheels, and martial arts remains challenging for current video diffusion models. Text-only conditioning is temporally ambiguous for fine-grained motion control, while explicit…

计算机视觉与模式识别 · 计算机科学 2026-03-10 Ashkan Taghipour , Morteza Ghahremani , Zinuo Li , Hamid Laga , Farid Boussaid , Mohammed Bennamoun

Unlike existing methods that rely on source images as appearance references and use source speech to generate motion, this work proposes a novel approach that directly extracts information from the speech, addressing key challenges in…

音频与语音处理 · 电气工程与系统科学 2026-03-03 Jinting Wang , Jun Wang , Hei Victor Cheng , Li Liu

While accurate lip synchronization has been achieved for arbitrary-subject audio-driven talking face generation, the problem of how to efficiently drive the head pose remains. Previous methods rely on pre-estimated structural information…

计算机视觉与模式识别 · 计算机科学 2021-04-23 Hang Zhou , Yasheng Sun , Wayne Wu , Chen Change Loy , Xiaogang Wang , Ziwei Liu

In this work, we address the problem of 4D facial expressions generation. This is usually addressed by animating a neutral 3D face to reach an expression peak, and then get back to the neutral state. In the real world though, people show…

计算机视觉与模式识别 · 计算机科学 2023-05-19 Naima Otberdout , Claudio Ferrari , Mohamed Daoudi , Stefano Berretti , Alberto Del Bimbo

Creating personalized 3D animations with precise control and realistic head motions remains challenging for current speech-driven 3D facial animation methods. Editing these animations is especially complex and time consuming, requires…

图形学 · 计算机科学 2025-10-01 Balamurugan Thambiraja , Malte Prinzler , Sadegh Aliakbarian , Darren Cosker , Justus Thies

Generating realistic human videos remains a challenging task, with the most effective methods currently relying on a human motion sequence as a control signal. Existing approaches often use existing motion extracted from other videos, which…

计算机视觉与模式识别 · 计算机科学 2024-12-18 Hsin-Ping Huang , Yang Zhou , Jui-Hsien Wang , Difan Liu , Feng Liu , Ming-Hsuan Yang , Zhan Xu

Audio-driven talking-head generation is a crucial and useful technology for virtual human interaction and film-making. While recent advances have focused on improving image fidelity and lip synchronization, generating accurate emotional…

计算机视觉与模式识别 · 计算机科学 2025-05-05 Wenqing Wang , Yun Fu

We present a novel approach for synthesizing 3D facial motions from audio sequences using key motion embeddings. Despite recent advancements in data-driven techniques, accurately mapping between audio signals and 3D facial meshes remains…

计算机视觉与模式识别 · 计算机科学 2024-09-04 Zhihao Xu , Shengjie Gong , Jiapeng Tang , Lingyu Liang , Yining Huang , Haojie Li , Shuangping Huang

Text-to-motion generation is an emerging and challenging problem, which aims to synthesize motion with the same semantics as the input text. However, due to the lack of diverse labeled training data, most approaches either limit to specific…

计算机视觉与模式识别 · 计算机科学 2023-03-27 Junfan Lin , Jianlong Chang , Lingbo Liu , Guanbin Li , Liang Lin , Qi Tian , Chang Wen Chen

Human-human communication is like a delicate dance where listeners and speakers concurrently interact to maintain conversational dynamics. Hence, an effective model for generating listener nonverbal behaviors requires understanding the…

计算机视觉与模式识别 · 计算机科学 2024-07-19 Minh Tran , Di Chang , Maksim Siniukov , Mohammad Soleymani

Human video generation is a dynamic and rapidly evolving task that aims to synthesize 2D human body video sequences with generative models given control conditions such as text, audio, and pose. With the potential for wide-ranging…

计算机视觉与模式识别 · 计算机科学 2024-07-12 Wentao Lei , Jinting Wang , Fengji Ma , Guanjie Huang , Li Liu

The goal of this work is to simultaneously generate natural talking faces and speech outputs from text. We achieve this by integrating Talking Face Generation (TFG) and Text-to-Speech (TTS) systems into a unified framework. We address the…

计算机视觉与模式识别 · 计算机科学 2024-05-17 Youngjoon Jang , Ji-Hoon Kim , Junseok Ahn , Doyeop Kwak , Hong-Sun Yang , Yoon-Cheol Ju , Il-Hwan Kim , Byeong-Yeol Kim , Joon Son Chung

Current co-speech motion generation approaches usually focus on upper body gestures following speech contents only, while lacking supporting the elaborate control of synergistic full-body motion based on text prompts, such as talking while…

计算机视觉与模式识别 · 计算机科学 2024-10-02 Bohong Chen , Yumeng Li , Yao-Xiang Ding , Tianjia Shao , Kun Zhou

Generating gestures from human speech has gained tremendous progress in animating virtual avatars. While the existing methods enable synthesizing gestures cooperated by individual self-talking, they overlook the practicality of concurrent…

计算机视觉与模式识别 · 计算机科学 2025-05-06 Xingqun Qi , Yatian Wang , Hengyuan Zhang , Jiahao Pan , Wei Xue , Shanghang Zhang , Wenhan Luo , Qifeng Liu , Yike Guo

This paper presents a generic method for generating full facial 3D animation from speech. Existing approaches to audio-driven facial animation exhibit uncanny or static upper face animation, fail to produce accurate and plausible…

计算机视觉与模式识别 · 计算机科学 2022-05-23 Alexander Richard , Michael Zollhoefer , Yandong Wen , Fernando de la Torre , Yaser Sheikh

To the best of our knowledge, we first present a live system that generates personalized photorealistic talking-head animation only driven by audio signals at over 30 fps. Our system contains three stages. The first stage is a deep neural…

图形学 · 计算机科学 2021-09-27 Yuanxun Lu , Jinxiang Chai , Xun Cao

In this paper we propose a convolutional autoencoder to address the problem of motion infilling for 3D human motion data. Given a start and end sequence, motion infilling aims to complete the missing gap in between, such that the filled in…

计算机视觉与模式识别 · 计算机科学 2021-11-17 Manuel Kaufmann , Emre Aksan , Jie Song , Fabrizio Pece , Remo Ziegler , Otmar Hilliges