English
Related papers

Related papers: X-Dancer: Expressive Music to Human Dance Video Ge…

200 papers

We present X2Video, the first diffusion model for rendering photorealistic videos guided by intrinsic channels including albedo, normal, roughness, metallicity, and irradiance, while supporting intuitive multi-modal controls with reference…

Graphics · Computer Science 2025-10-10 Zhitong Huang , Mohan Zhang , Renhan Wang , Rui Tang , Hao Zhu , Jing Liao

Generating realistic reactive motions, in which one person reacts to the fixed motions of others, is challenging due to strict interaction constraints and a limited feasible solution space. This paper focuses on a typical scenario: duet…

Computer Vision and Pattern Recognition · Computer Science 2026-05-07 Xuhai Chen , Zhi Cen , Huaijin Pi , Sida Peng , Xiaowei Zhou , Yong Liu

Dance plays an important role as an artistic form and expression in human culture, yet automatically generating dance sequences is a significant yet challenging endeavor. Existing approaches often neglect the critical aspect of…

Computer Vision and Pattern Recognition · Computer Science 2026-03-11 Hongsong Wang , Ying Zhu , Xin Geng , Liang Wang

Diffusion models usher a new era of video editing, flexibly manipulating the video contents with text prompts. Despite the widespread application demand in editing human-centered videos, these models face significant challenges in handling…

Computer Vision and Pattern Recognition · Computer Science 2024-08-15 Xiaojing Zhong , Xinyi Huang , Xiaofeng Yang , Guosheng Lin , Qingyao Wu

Dance is a form of human motion characterized by emotional expression and communication, playing a role in various fields such as music, virtual reality, and content creation. Existing methods for dance generation often fail to adequately…

Computer Vision and Pattern Recognition · Computer Science 2026-03-10 Sangjune Park , Inhyeok Choi , Donghyeon Soon , Youngwoo Jeon , Kyungdon Joo

Lip-syncing videos with given audio is the foundation for various applications including the creation of virtual presenters or performers. While recent studies explore high-fidelity lip-sync with different techniques, their task-orientated…

Computer Vision and Pattern Recognition · Computer Science 2024-08-07 Jiazhi Guan , Zhiliang Xu , Hang Zhou , Kaisiyuan Wang , Shengyi He , Zhanwang Zhang , Borong Liang , Haocheng Feng , Errui Ding , Jingtuo Liu , Jingdong Wang , Youjian Zhao , Ziwei Liu

We present AnimaX, a feed-forward 3D animation framework that bridges the motion priors of video diffusion models with the controllable structure of skeleton-based animation. Traditional motion synthesis methods are either restricted to…

Computer Vision and Pattern Recognition · Computer Science 2025-06-25 Zehuan Huang , Haoran Feng , Yangtian Sun , Yuanchen Guo , Yanpei Cao , Lu Sheng

Generating animation of physics-based characters with intuitive control has long been a desirable task with numerous applications. However, generating physically simulated animations that reflect high-level human instructions remains a…

Computer Vision and Pattern Recognition · Computer Science 2023-12-29 Jiawei Ren , Mingyuan Zhang , Cunjun Yu , Xiao Ma , Liang Pan , Ziwei Liu

Human motion generation aims to produce plausible human motion sequences according to various conditional inputs, such as text or audio. Despite the feasibility of existing methods in generating motion based on short prompts and simple…

Multimedia · Computer Science 2024-11-12 Bo Han , Hao Peng , Minjing Dong , Yi Ren , Yixuan Shen , Chang Xu

Human-centric volumetric videos offer immersive free-viewpoint experiences, yet existing methods focus either on replaying general dynamic scenes or animating human avatars, limiting their ability to re-perform general dynamic scenes. In…

Computer Vision and Pattern Recognition · Computer Science 2025-03-18 Yuheng Jiang , Zhehao Shen , Chengcheng Guo , Yu Hong , Zhuo Su , Yingliang Zhang , Marc Habermann , Lan Xu

Existing methods for image-to-3D avatar generation struggle to produce highly detailed, animation-ready avatars suitable for real-world applications. We introduce AdaHuman, a novel framework that generates high-fidelity animatable 3D…

Computer Vision and Pattern Recognition · Computer Science 2025-06-02 Yangyi Huang , Ye Yuan , Xueting Li , Jan Kautz , Umar Iqbal

Speech-driven gesture generation is an emerging field within virtual human creation. However, a significant challenge lies in accurately determining and processing the multitude of input features (such as acoustic, semantic, emotional,…

Sound · Computer Science 2024-03-19 Fan Zhang , Zhaohan Wang , Xin Lyu , Siyuan Zhao , Mengjian Li , Weidong Geng , Naye Ji , Hui Du , Fuxing Gao , Hao Wu , Shunman Li

Human video synthesis aims to create lifelike characters in various environments, with wide applications in VR, storytelling, and content creation. While 2D diffusion-based methods have made significant progress, they struggle to generalize…

Computer Vision and Pattern Recognition · Computer Science 2024-12-19 Liyuan Cui , Xiaogang Xu , Wenqi Dong , Zesong Yang , Hujun Bao , Zhaopeng Cui

Motion-to-music and music-to-motion have been studied separately, each attracting substantial research interest within their respective domains. The interaction between human motion and music is a reflection of advanced human intelligence,…

Sound · Computer Science 2024-11-05 Fuming You , Minghui Fang , Li Tang , Rongjie Huang , Yongqi Wang , Zhou Zhao

Human video generation is a dynamic and rapidly evolving task that aims to synthesize 2D human body video sequences with generative models given control conditions such as text, audio, and pose. With the potential for wide-ranging…

Computer Vision and Pattern Recognition · Computer Science 2024-07-12 Wentao Lei , Jinting Wang , Fengji Ma , Guanjie Huang , Li Liu

We propose DiffSHEG, a Diffusion-based approach for Speech-driven Holistic 3D Expression and Gesture generation with arbitrary length. While previous works focused on co-speech gesture or expression generation individually, the joint…

Sound · Computer Science 2024-04-09 Junming Chen , Yunfei Liu , Jianan Wang , Ailing Zeng , Yu Li , Qifeng Chen

Video generation is an inherently challenging task, as it requires modeling realistic temporal dynamics as well as spatial content. Existing methods entangle the two intrinsically different tasks of motion and content creation in a single…

Computer Vision and Pattern Recognition · Computer Science 2020-01-13 Ximeng Sun , Huijuan Xu , Kate Saenko

In this study, we introduce a learning-based method for generating high-quality human motion sequences from text descriptions (e.g., ``A person walks forward"). Existing techniques struggle with motion diversity and smooth transitions in…

Computer Vision and Pattern Recognition · Computer Science 2023-12-08 Weilin Wan , Yiming Huang , Shutong Wu , Taku Komura , Wenping Wang , Dinesh Jayaraman , Lingjie Liu

Text-driven human motion generation is an emerging task in animation and humanoid robot design. Existing algorithms directly generate the full sequence which is computationally expensive and prone to errors as it does not pay special…

Computer Vision and Pattern Recognition · Computer Science 2024-05-27 Zichen Geng , Caren Han , Zeeshan Hayder , Jian Liu , Mubarak Shah , Ajmal Mian

In recent years, video generation has seen significant advancements. However, challenges still persist in generating complex motions and interactions. To address these challenges, we introduce ReVision, a plug-and-play framework that…

Computer Vision and Pattern Recognition · Computer Science 2026-01-12 Qihao Liu , Ju He , Qihang Yu , Liang-Chieh Chen , Alan Yuille
‹ Prev 1 8 9 10 Next ›