中文
相关论文

相关论文: Emotional Speech-driven 3D Body Animation via Dise…

200 篇论文

Recent talking avatar generation models have made strides in achieving realistic and accurate lip synchronization with the audio, but often fall short in controlling and conveying detailed expressions and emotions of the avatar, making the…

计算机视觉与模式识别 · 计算机科学 2024-05-27 Yuchi Wang , Junliang Guo , Jianhong Bai , Runyi Yu , Tianyu He , Xu Tan , Xu Sun , Jiang Bian

Talking head synthesis is a promising approach for the video production industry. Recently, a lot of effort has been devoted in this research area to improve the generation quality or enhance the model generalization. However, there are few…

计算机视觉与模式识别 · 计算机科学 2023-04-21 Shuai Shen , Wenliang Zhao , Zibin Meng , Wanhua Li , Zheng Zhu , Jie Zhou , Jiwen Lu

Human speech is often accompanied by hand and arm gestures. Given audio speech input, we generate plausible gestures to go along with the sound. Specifically, we perform cross-modal translation from "in-the-wild'' monologue speech of a…

计算机视觉与模式识别 · 计算机科学 2019-06-11 Shiry Ginosar , Amir Bar , Gefen Kohavi , Caroline Chan , Andrew Owens , Jitendra Malik

Pose-driven full-body avatars built on neural rendering produce high-quality novel views of a captured subject. Yet loose clothing and other dynamic elements deform in ways pose alone cannot explain: the same pose can correspond to many…

计算机视觉与模式识别 · 计算机科学 2026-05-21 Shichong Peng , Chengxiang Yin , Fei Jiang , Zhongshi Jiang , Lingchen Yang , Qingyang Tan , Amin Jourabloo , Jason Saragih , Ke Li , Christian Häne

We introduce SEDTalker, an emotion-aware framework for speech-driven 3D facial animation that leverages frame-level speech emotion diarization to achieve fine-grained expressive control. Unlike prior approaches that rely on utterance-level…

计算机视觉与模式识别 · 计算机科学 2026-04-16 Farzaneh Jafari , Stefano Berretti , Anup Basu

Previous motion generation methods are limited to the pre-rigged 3D human model, hindering their applications in the animation of various non-rigged characters. In this work, we present TapMo, a Text-driven Animation Pipeline for…

图形学 · 计算机科学 2023-10-20 Jiaxu Zhang , Shaoli Huang , Zhigang Tu , Xin Chen , Xiaohang Zhan , Gang Yu , Ying Shan

We introduce FactorPortrait, a video diffusion method for controllable portrait animation that enables lifelike synthesis from disentangled control signals of facial expressions, head movement, and camera viewpoints. Given a single portrait…

计算机视觉与模式识别 · 计算机科学 2025-12-15 Jiapeng Tang , Kai Li , Chengxiang Yin , Liuhao Ge , Fei Jiang , Jiu Xu , Matthias Nießner , Christian Häne , Timur Bagautdinov , Egor Zakharov , Peihong Guo

We present a generative adversarial network to synthesize 3D pose sequences of co-speech upper-body gestures with appropriate affective expressions. Our network consists of two components: a generator to synthesize gestures from a joint…

多媒体 · 计算机科学 2024-11-26 Uttaran Bhattacharya , Elizabeth Childs , Nicholas Rewkowski , Dinesh Manocha

Speech-driven talking head generation is a critical yet challenging task with applications in augmented reality and virtual human modeling. While recent approaches using autoregressive and diffusion-based models have achieved notable…

计算机视觉与模式识别 · 计算机科学 2025-12-08 Yihong Lin , Zhaoxin Fan , Xianjia Wu , Lingyu Xiong , Liang Peng , Xiandong Li , Wenxiong Kang , Songju Lei , Huang Xu

Co-speech gesture generation is crucial for producing synchronized and realistic human gestures that accompany speech, enhancing the animation of lifelike avatars in virtual environments. While diffusion models have shown impressive…

Recently, diffusion models have shown their impressive ability in visual generation tasks. Besides static images, more and more research attentions have been drawn to the generation of realistic videos. The video generation not only has a…

计算机视觉与模式识别 · 计算机科学 2025-11-03 Yucheng Xing , Jinxing Yin , Xiaodong Liu

Diffusion models can synthesize realistic co-speech video from audio for various applications, such as video creation and virtual agents. However, existing diffusion-based methods are slow due to numerous denoising steps and costly…

计算机视觉与模式识别 · 计算机科学 2025-10-06 Beijia Lu , Ziyi Chen , Jing Xiao , Jun-Yan Zhu

Diffusion models have shown great promise for image generation, beating GANs in terms of generation diversity, with comparable image quality. However, their application to 3D shapes has been limited to point or voxel representations that…

计算机视觉与模式识别 · 计算机科学 2022-12-16 Gimin Nam , Mariem Khlifi , Andrew Rodriguez , Alberto Tono , Linqi Zhou , Paul Guerrero

We present aMUSEd, an open-source, lightweight masked image model (MIM) for text-to-image generation based on MUSE. With 10 percent of MUSE's parameters, aMUSEd is focused on fast image generation. We believe MIM is under-explored compared…

计算机视觉与模式识别 · 计算机科学 2024-01-04 Suraj Patil , William Berman , Robin Rombach , Patrick von Platen

Generating natural, correct, and visually smooth 3D avatar sign language motion conditioned on the text inputs continues to be very challenging. In this work, we train a generative model of 3D body motion and explore the role of…

计算机视觉与模式识别 · 计算机科学 2026-03-31 Rui Hong , Jana Kosecka

Controllable generation of 3D human motions becomes an important topic as the world embraces digital transformation. Existing works, though making promising progress with the advent of diffusion models, heavily rely on meticulously captured…

计算机视觉与模式识别 · 计算机科学 2024-01-25 Nhat M. Hoang , Kehong Gong , Chuan Guo , Michael Bi Mi

Generating realistic conversational gestures are essential for achieving natural, socially engaging interactions with digital humans. However, existing methods typically map a single audio stream to a single speaker's motion, without…

计算机视觉与模式识别 · 计算机科学 2026-03-31 Yichen Peng , Jyun-Ting Song , Siyeol Jung , Ruofan Liu , Haiyang Liu , Xuangeng Chu , Ruicong Liu , Erwin Wu , Hideki Koike , Kris Kitani

In this paper, we introduce GaussianMotion, a novel human rendering model that generates fully animatable scenes aligned with textual descriptions using Gaussian Splatting. Although existing methods achieve reasonable text-to-3D generation…

计算机视觉与模式识别 · 计算机科学 2025-02-18 Gyumin Shim , Sangmin Lee , Jaegul Choo

With the introduction of diffusion-based video generation techniques, audio-conditioned human video generation has recently achieved significant breakthroughs in both the naturalness of motion and the synthesis of portrait details. Due to…

计算机视觉与模式识别 · 计算机科学 2025-04-07 Jianwen Jiang , Chao Liang , Jiaqi Yang , Gaojie Lin , Tianyun Zhong , Yanbo Zheng

We introduce AvatarBooth, a novel method for generating high-quality 3D avatars using text prompts or specific images. Unlike previous approaches that can only synthesize avatars based on simple text descriptions, our method enables the…

计算机视觉与模式识别 · 计算机科学 2023-06-19 Yifei Zeng , Yuanxun Lu , Xinya Ji , Yao Yao , Hao Zhu , Xun Cao
‹ 上一页 1 8 9 10 下一页 ›