中文
相关论文

相关论文: Learning Speech-driven 3D Conversational Gestures …

200 篇论文

3D Gaussian splatting-based talking head synthesis has recently gained attention for its ability to render high-fidelity images with real-time inference speed. However, since it is typically trained on only a short video that lacks the…

计算机视觉与模式识别 · 计算机科学 2025-02-04 Junuk Cha , Seongro Yoon , Valeriya Strizhkova , Francois Bremond , Seungryul Baek

Technological developments have produced methods that can generate educational videos from input text or sound. Recently, the use of deep learning techniques for image and video generation has been widely explored, particularly in…

多媒体 · 计算机科学 2026-01-27 M. E. ElAlami , S. M. Khater , M. El. R. Rehan

Generating holistic co-speech gestures that integrate full-body motion with facial expressions suffers from semantically incoherent coordination on body motion and spatially unstable meaningless movements due to existing part-decomposed or…

计算机视觉与模式识别 · 计算机科学 2026-01-27 Xuanmeng Sha , Liyun Zhang , Tomohiro Mashita , Naoya Chiba , Yuki Uranishi

Novel-View Human Action Synthesis aims to synthesize the movement of a body from a virtual viewpoint, given a video from a real viewpoint. We present a novel 3D reasoning to synthesize the target viewpoint. We first estimate the 3D mesh of…

计算机视觉与模式识别 · 计算机科学 2020-10-09 Mohamed Ilyes Lakhal , Davide Boscaini , Fabio Poiesi , Oswald Lanz , Andrea Cavallaro

Animation of humanoid characters is essential in various graphics applications, but requires significant time and cost to create realistic animations. We propose an approach to synthesize 4D animated sequences of input static 3D humanoid…

图形学 · 计算机科学 2025-03-21 Marc Benedí San Millán , Angela Dai , Matthias Nießner

We present a novel approach for synthesizing 3D talking heads with controllable emotion, featuring enhanced lip synchronization and rendering quality. Despite significant progress in the field, prior methods still suffer from multi-view…

计算机视觉与模式识别 · 计算机科学 2024-08-02 Qianyun He , Xinya Ji , Yicheng Gong , Yuanxun Lu , Zhengyu Diao , Linjia Huang , Yao Yao , Siyu Zhu , Zhan Ma , Songcen Xu , Xiaofei Wu , Zixiao Zhang , Xun Cao , Hao Zhu

Generating videos of complex human motions such as flips, cartwheels, and martial arts remains challenging for current video diffusion models. Text-only conditioning is temporally ambiguous for fine-grained motion control, while explicit…

计算机视觉与模式识别 · 计算机科学 2026-03-10 Ashkan Taghipour , Morteza Ghahremani , Zinuo Li , Hamid Laga , Farid Boussaid , Mohammed Bennamoun

Hand pose estimation from a monocular RGB image is an important but challenging task. The main factor affecting its performance is the lack of a sufficiently large training dataset with accurate hand-keypoint annotations. In this work, we…

计算机视觉与模式识别 · 计算机科学 2020-12-15 Liangjian Chen , Shih-Yao Lin , Yusheng Xie , Hui Tang , Yufan Xue , Xiaohui Xie , Yen-Yu Lin , Wei Fan

Synthesizing dynamic appearances of humans in motion plays a central role in applications such as AR/VR and video editing. While many recent methods have been proposed to tackle this problem, handling loose garments with complex textures…

计算机视觉与模式识别 · 计算机科学 2021-11-12 Tuanfeng Y. Wang , Duygu Ceylan , Krishna Kumar Singh , Niloy J. Mitra

We present an audio-driven real-time system for animating photorealistic 3D facial avatars with minimal latency, designed for social interactions in virtual reality for anyone. Central to our approach is an encoder model that transforms…

图形学 · 计算机科学 2025-11-04 Jiye Lee , Chenghui Li , Linh Tran , Shih-En Wei , Jason Saragih , Alexander Richard , Hanbyul Joo , Shaojie Bai

Speechreading is a notoriously difficult task for humans to perform. In this paper we present an end-to-end model based on a convolutional neural network (CNN) for generating an intelligible acoustic speech signal from silent video frames…

计算机视觉与模式识别 · 计算机科学 2017-01-10 Ariel Ephrat , Shmuel Peleg

We propose DiffSHEG, a Diffusion-based approach for Speech-driven Holistic 3D Expression and Gesture generation with arbitrary length. While previous works focused on co-speech gesture or expression generation individually, the joint…

声音 · 计算机科学 2024-04-09 Junming Chen , Yunfei Liu , Jianan Wang , Ailing Zeng , Yu Li , Qifeng Chen

Synthesizing realistic videos of humans using neural networks has been a popular alternative to the conventional graphics-based rendering pipeline due to its high efficiency. Existing works typically formulate this as an image-to-image…

Both acoustic and visual information influence human perception of speech. For this reason, the lack of audio in a video sequence determines an extremely low speech intelligibility for untrained lip readers. In this paper, we present a way…

音频与语音处理 · 电气工程与系统科学 2020-08-18 Daniel Michelsanti , Olga Slizovskaia , Gloria Haro , Emilia Gómez , Zheng-Hua Tan , Jesper Jensen

The creation of virtual humans increasingly leverages automated synthesis of speech and gestures, enabling expressive, adaptable agents that effectively engage users. However, the independent development of voice and gesture generation…

图形学 · 计算机科学 2025-07-02 Haoyang Du , Kiran Chhatre , Christopher Peters , Brian Keegan , Rachel McDonnell , Cathy Ennis

Generating coherent and useful image/video scenes from a free-form textual description is technically a very difficult problem to handle. Textual description of the same scene can vary greatly from person to person, or sometimes even for…

计算机视觉与模式识别 · 计算机科学 2020-12-01 Faria Huq , Nafees Ahmed , Anindya Iqbal

Natural co-speech gestures are essential components to improve the experience of Human-robot interaction (HRI). However, current gesture generation approaches have many limitations of not being natural, not aligning with the speech and…

机器人学 · 计算机科学 2024-09-10 Shiyi Tang , Christian Dondrup

Gesture synthesis has gained significant attention as a critical research field, aiming to produce contextually appropriate and natural gestures corresponding to speech or textual input. Although deep learning-based approaches have achieved…

计算与语言 · 计算机科学 2024-05-29 Nan Gao , Zeyu Zhao , Zhi Zeng , Shuwu Zhang , Dongdong Weng , Yihua Bao

Fast and robust three-dimensional reconstruction of facial geometric structure from a single image is a challenging task with numerous applications. Here, we introduce a learning-based approach for reconstructing a three-dimensional face…

计算机视觉与模式识别 · 计算机科学 2016-09-27 Elad Richardson , Matan Sela , Ron Kimmel

Audio-driven 3D facial animation has several virtual humans applications for content creation and editing. While several existing methods provide solutions for speech-driven animation, precise control over content (what) and style (how) of…

声音 · 计算机科学 2024-08-15 Qingju Liu , Hyeongwoo Kim , Gaurav Bharaj
‹ 上一页 1 8 9 10 下一页 ›