English
Related papers

Related papers: MagicPose4D: Crafting Articulated Models with Appe…

200 papers

Video world models can generate realistic futures from a single instruction, but they often fail to preserve consistent point-level motion over time. As a result, the generated videos appear plausible, yet lack the physical grounding…

Computer Vision and Pattern Recognition · Computer Science 2026-05-25 Kaichen Zhou , Yuzhen Chen , Fangneng Zhan , Hang Hua , Grace Chen , Xinhai Chang , Ao Qu , Yilun Du , Zhuang Liu , Paul Pu Liang , Mengyu Wang

Benefiting from the rapid development of 2D diffusion models, 3D content generation has witnessed significant progress. One promising solution is to finetune the pre-trained 2D diffusion models to produce multi-view images and then…

Computer Vision and Pattern Recognition · Computer Science 2025-01-10 Fan Yang , Jianfeng Zhang , Yichun Shi , Bowen Chen , Chenxu Zhang , Huichao Zhang , Xiaofeng Yang , Xiu Li , Jiashi Feng , Guosheng Lin

Recent pose-to-video models can translate 2D pose sequences into photorealistic, identity-preserving dance videos, so the key challenge is to generate temporally coherent, rhythm-aligned 2D poses from music, especially under complex,…

Computer Vision and Pattern Recognition · Computer Science 2025-12-15 Yan Zhang , Han Zou , Lincong Feng , Cong Xie , Ruiqi Yu , Zhenpeng Zhan

Reconstructing the motion of objects from videos is a key component for embodied AI and robot manipulation. While diverse approaches to object pose tracking have been studied, they rely heavily on strong external priors, such as depth data…

Computer Vision and Pattern Recognition · Computer Science 2026-05-25 Jisu Shin , Junoh Lee , JunGyu Lee , Inhwan Bae , Dohyeon Lee , Hokyun Im , Youngwoon Lee , Hae-Gon Jeon

Motion-based controllable video generation offers the potential for creating captivating visual content. Existing methods typically necessitate model training to encode particular motion cues or incorporate fine-tuning to inject certain…

Computer Vision and Pattern Recognition · Computer Science 2024-10-23 Pengyang Ling , Jiazi Bu , Pan Zhang , Xiaoyi Dong , Yuhang Zang , Tong Wu , Huaian Chen , Jiaqi Wang , Yi Jin

In this paper, we propose a novel learning approach for feed-forward one-shot 4D head avatar synthesis. Different from existing methods that often learn from reconstructing monocular videos guided by 3DMM, we employ pseudo multi-view videos…

Computer Vision and Pattern Recognition · Computer Science 2024-07-12 Yu Deng , Duomin Wang , Baoyuan Wang

Understanding the 3D motion of articulated objects is essential in robotic scene understanding, mobile manipulation, and motion planning. Prior methods for articulation estimation have primarily focused on controlled settings, assuming…

World models that support controllable and editable spatiotemporal environments are valuable for robotics, enabling scalable training data, repro ducible evaluation, and flexible task design. While recent text-to-video models generate…

Computer Vision and Pattern Recognition · Computer Science 2025-10-07 Xuehai He , Shijie Zhou , Thivyanth Venkateswaran , Kaizhi Zheng , Ziyu Wan , Achuta Kadambi , Xin Eric Wang

Dense 3D reconstruction and tracking of dynamic scenes from monocular video remains an important open challenge in computer vision. Progress in this area has been constrained by the scarcity of high-quality datasets with dense, complete,…

Computer Vision and Pattern Recognition · Computer Science 2026-05-07 Zeren Jiang , Yushi Lan , Yihang Luo , Yufan Deng , Zihang Lai , Edgar Sucar , Christian Rupprecht , Iro Laina , Diane Larlus , Chuanxia Zheng , Andrea Vedaldi

Musicians delicately control their bodies to generate music. Sometimes, their motions are too subtle to be captured by the human eye. To analyze how they move to produce the music, we need to estimate precise 4D human pose (3D pose over…

Computer Vision and Pattern Recognition · Computer Science 2024-11-27 Seong Jong Yoo , Snehesh Shrestha , Irina Muresanu , Cornelia Fermüller

One-shot video-driven talking face generation aims at producing a synthetic talking video by transferring the facial motion from a video to an arbitrary portrait image. Head pose and facial expression are always entangled in facial motion…

Computer Vision and Pattern Recognition · Computer Science 2023-03-03 Youxin Pang , Yong Zhang , Weize Quan , Yanbo Fan , Xiaodong Cun , Ying Shan , Dong-ming Yan

Four-dimensional MRI (4D-MRI) is an promising technique for capturing respiratory-induced motion in radiation therapy planning and delivery. Conventional 4D reconstruction methods, which typically rely on phase binning or separate template…

Computer Vision and Pattern Recognition · Computer Science 2025-09-24 Xinyang Wu , Muheng Li , Xia Li , Orso Pusterla , Sairos Safai , Philippe C. Cattin , Antony J. Lomax , Ye Zhang

Editing 4D scenes reconstructed from monocular videos based on text prompts is a valuable yet challenging task with broad applications in content creation and virtual environments. The key difficulty lies in achieving semantically precise…

Computer Vision and Pattern Recognition · Computer Science 2025-10-13 Jin-Chuan Shi , Chengye Su , Jiajun Wang , Ariel Shamir , Miao Wang

This paper presents a unified approach to understanding dynamic scenes from casual videos. Large pretrained vision foundation models, such as vision-language, video depth prediction, motion tracking, and segmentation models, offer promising…

Computer Vision and Pattern Recognition · Computer Science 2025-03-28 David Yifan Yao , Albert J. Zhai , Shenlong Wang

We present I2V3D, a novel framework for animating static images into dynamic videos with precise 3D control, leveraging the strengths of both 3D geometry guidance and advanced generative models. Our approach combines the precision of a…

Computer Vision and Pattern Recognition · Computer Science 2025-03-14 Zhiyuan Zhang , Dongdong Chen , Jing Liao

We present HaPTIC, an approach that infers coherent 4D hand trajectories from monocular videos. Current video-based hand pose reconstruction methods primarily focus on improving frame-wise 3D pose using adjacent frames rather than studying…

Computer Vision and Pattern Recognition · Computer Science 2025-01-15 Yufei Ye , Yao Feng , Omid Taheri , Haiwen Feng , Shubham Tulsiani , Michael J. Black

Generating text-editable and pose-controllable character videos have an imperious demand in creating various digital human. Nevertheless, this task has been restricted by the absence of a comprehensive dataset featuring paired video-pose…

Computer Vision and Pattern Recognition · Computer Science 2024-01-04 Yue Ma , Yingqing He , Xiaodong Cun , Xintao Wang , Siran Chen , Ying Shan , Xiu Li , Qifeng Chen

Text-to-motion (T2M) generation is becoming a practical tool for animation and interactive avatars. However, modifying specific body parts while maintaining overall motion coherence remains challenging. Existing methods typically rely on…

Computer Vision and Pattern Recognition · Computer Science 2026-03-23 Minyue Dai , Ke Fan , Anyi Rao , Jingbo Wang , Bo Dai

We introduce UPose3D, a novel approach for multi-view 3D human pose estimation, addressing challenges in accuracy and scalability. Our method advances existing pose estimation frameworks by improving robustness and flexibility without…

Computer Vision and Pattern Recognition · Computer Science 2024-07-11 Vandad Davoodnia , Saeed Ghorbani , Marc-André Carbonneau , Alexandre Messier , Ali Etemad

World models aim to endow AI systems with the ability to represent, generate, and interact with dynamic environments in a coherent and temporally consistent manner. While recent video generation models have demonstrated impressive visual…