English
Related papers

Related papers: Motion-X++: A Large-Scale Multimodal 3D Whole-body…

200 papers

Recent advances in diffusion models have significantly improved conditional video generation, particularly in the pose-guided human image animation task. Although existing methods are capable of generating high-fidelity and time-consistent…

Computer Vision and Pattern Recognition · Computer Science 2026-03-19 Shuolin Xu , Siming Zheng , Ziyi Wang , HC Yu , Jinwei Chen , Huaqi Zhang , Daquan Zhou , Tong-Yee Lee , Bo Li , Peng-Tao Jiang

In filmmaking, directors typically allow actors to perform freely based on the script before providing specific guidance on how to present key actions. AI-generated content faces similar requirements, where users not only need automatic…

Computer Vision and Pattern Recognition · Computer Science 2025-07-29 Zheng Qin , Ruobing Zheng , Yabing Wang , Tianqi Li , Zixin Zhu , Sanping Zhou , Ming Yang , Le Wang

This work addresses the problem of generating 3D holistic body motions from human speech. Given a speech recording, we synthesize sequences of 3D body poses, hand gestures, and facial expressions that are realistic and diverse. To achieve…

Computer Vision and Pattern Recognition · Computer Science 2023-06-21 Hongwei Yi , Hualin Liang , Yifei Liu , Qiong Cao , Yandong Wen , Timo Bolkart , Dacheng Tao , Michael J. Black

We introduce X-Dyna, a novel zero-shot, diffusion-based pipeline for animating a single human image using facial expressions and body movements derived from a driving video, that generates realistic, context-aware dynamics for both the…

Computer Vision and Pattern Recognition · Computer Science 2025-01-22 Di Chang , Hongyi Xu , You Xie , Yipeng Gao , Zhengfei Kuang , Shengqu Cai , Chenxu Zhang , Guoxian Song , Chao Wang , Yichun Shi , Zeyuan Chen , Shijie Zhou , Linjie Luo , Gordon Wetzstein , Mohammad Soleymani

In this paper, we tackle the problem of how to build and benchmark a large motion model (LMM). The ultimate goal of LMM is to serve as a foundation model for versatile motion-related tasks, e.g., human motion generation, with…

Computer Vision and Pattern Recognition · Computer Science 2024-10-18 Liang Xu , Shaoyang Hua , Zili Lin , Yifan Liu , Feipeng Ma , Yichao Yan , Xin Jin , Xiaokang Yang , Wenjun Zeng

We present DPoser-X, a diffusion-based prior model for 3D whole-body human poses. Building a versatile and robust full-body human pose prior remains challenging due to the inherent complexity of articulated human poses and the scarcity of…

Computer Vision and Pattern Recognition · Computer Science 2025-08-05 Junzhe Lu , Jing Lin , Hongkun Dou , Ailing Zeng , Yue Deng , Xian Liu , Zhongang Cai , Lei Yang , Yulun Zhang , Haoqian Wang , Ziwei Liu

This paper proposes a novel application system for the generation of three-dimensional (3D) character animation driven by markerless human body motion capturing. The entire pipeline of the system consists of five stages: 1) the capturing of…

Computer Vision and Pattern Recognition · Computer Science 2022-12-13 Jinbao Wang , Ke Lu , Jian Xue

We present X-Avatar, a novel avatar model that captures the full expressiveness of digital humans to bring about life-like experiences in telepresence, AR/VR and beyond. Our method models bodies, hands, facial expressions and appearance in…

Computer Vision and Pattern Recognition · Computer Science 2023-03-10 Kaiyue Shen , Chen Guo , Manuel Kaufmann , Juan Jose Zarate , Julien Valentin , Jie Song , Otmar Hilliges

Though the advancement of pre-trained large language models unfolds, the exploration of building a unified model for language and other multi-modal data, such as motion, remains challenging and untouched so far. Fortunately, human motion…

Computer Vision and Pattern Recognition · Computer Science 2023-07-21 Biao Jiang , Xin Chen , Wen Liu , Jingyi Yu , Gang Yu , Tao Chen

We introduce Unimotion, the first unified multi-task human motion model capable of both flexible motion control and frame-level motion understanding. While existing works control avatar motion with global text conditioning, or with…

Computer Vision and Pattern Recognition · Computer Science 2024-10-01 Chuqiao Li , Julian Chibane , Yannan He , Naama Pearl , Andreas Geiger , Gerard Pons-moll

Text-to-motion generation has experienced remarkable progress in recent years. However, current approaches remain limited to synthesizing motion from short or general text prompts, primarily due to dataset constraints. This limitation…

Computer Vision and Pattern Recognition · Computer Science 2025-10-24 Chuan Guo , Inwoo Hwang , Jian Wang , Bing Zhou

We present AIST++, a new multi-modal dataset of 3D dance motion and music, along with FACT, a Full-Attention Cross-modal Transformer network for generating 3D dance motion conditioned on music. The proposed AIST++ dataset contains 5.2 hours…

Computer Vision and Pattern Recognition · Computer Science 2021-08-03 Ruilong Li , Shan Yang , David A. Ross , Angjoo Kanazawa

We introduce HuMoR: a 3D Human Motion Model for Robust Estimation of temporal pose and shape. Though substantial progress has been made in estimating 3D human motion and shape from dynamic observations, recovering plausible pose sequences…

Computer Vision and Pattern Recognition · Computer Science 2021-08-19 Davis Rempe , Tolga Birdal , Aaron Hertzmann , Jimei Yang , Srinath Sridhar , Leonidas J. Guibas

Natural language plays a critical role in many computer vision applications, such as image captioning, visual question answering, and cross-modal retrieval, to provide fine-grained semantic information. Unfortunately, while human pose is…

Computer Vision and Pattern Recognition · Computer Science 2024-09-11 Ginger Delmas , Philippe Weinzaepfel , Thomas Lucas , Francesc Moreno-Noguer , Grégory Rogez

State-of-the-art text-to-motion generation models rely on the kinematic-aware, local-relative motion representation popularized by HumanML3D, which encodes motion relative to the pelvis and to the previous frame with built-in redundancy.…

Computer Vision and Pattern Recognition · Computer Science 2025-06-03 Zichong Meng , Zeyu Han , Xiaogang Peng , Yiming Xie , Huaizu Jiang

Despite the fact that many 3D human activity benchmarks being proposed, most existing action datasets focus on the action recognition tasks for the segmented videos. There is a lack of standard large-scale benchmarks, especially for current…

Computer Vision and Pattern Recognition · Computer Science 2017-03-29 Chunhui Liu , Yueyu Hu , Yanghao Li , Sijie Song , Jiaying Liu

This work presents 4DHumanOutfit, a new dataset of densely sampled spatio-temporal 4D human motion data of different actors, outfits and motions. The dataset is designed to contain different actors wearing different outfits while performing…

When executing whole-body motions, humans are able to use a large variety of support poses which not only utilize the feet, but also hands, knees and elbows to enhance stability. While there are many works analyzing the transitions involved…

Robotics · Computer Science 2015-10-01 Christian Mandery , Júlia Borràs , Mirjam Jöchner , Tamim Asfour

We present AnimaX, a feed-forward 3D animation framework that bridges the motion priors of video diffusion models with the controllable structure of skeleton-based animation. Traditional motion synthesis methods are either restricted to…

Computer Vision and Pattern Recognition · Computer Science 2025-06-25 Zehuan Huang , Haoran Feng , Yangtian Sun , Yuanchen Guo , Yanpei Cao , Lu Sheng

Monocular egocentric 3D human motion capture remains a significant challenge, particularly under conditions of low lighting and fast movements, which are common in head-mounted device applications. Existing methods that rely on RGB cameras…

Computer Vision and Pattern Recognition · Computer Science 2025-02-13 Christen Millerdurai , Hiroyasu Akada , Jian Wang , Diogo Luvizon , Alain Pagani , Didier Stricker , Christian Theobalt , Vladislav Golyanik