English
Related papers

Related papers: OrthoPhys: Physically Plausible Video Generation w…

200 papers

We present a generative method to estimate 3D human motion and body shape from monocular video. Under the assumption that starting from an initial pose optical flow constrains subsequent human motion, we exploit flow to find temporally…

Computer Vision and Pattern Recognition · Computer Science 2017-03-22 Thiemo Alldieck , Marc Kassubeck , Marcus Magnor

Learning to estimate 3D geometry in a single image by watching unlabeled videos via deep convolutional network has made significant process recently. Current state-of-the-art (SOTA) methods, are based on the learning framework of rigid…

Computer Vision and Pattern Recognition · Computer Science 2018-08-17 Zhenheng Yang , Peng Wang , Yang Wang , Wei Xu , Ram Nevatia

Video generation aims to produce temporally coherent sequences of visual frames, representing a pivotal advancement in Artificial Intelligence Generated Content (AIGC). Compared to static image generation, video generation poses unique…

Computer Vision and Pattern Recognition · Computer Science 2026-02-19 Zhiyu Yin , Kehai Chen , Xuefeng Bai , Ruili Jiang , Juntao Li , Hongdong Li , Jin Liu , Yang Xiang , Jun Yu , Min Zhang

Videos inherently represent 2D projections of a dynamic 3D world. However, our analysis suggests that video diffusion models trained solely on raw video data often fail to capture meaningful geometric-aware structure in their learned…

Computer Vision and Pattern Recognition · Computer Science 2026-05-06 Haoyu Wu , Diankun Wu , Tianyu He , Junliang Guo , Yang Ye , Yueqi Duan , Jiang Bian

Generative models have achieved success in producing apparently coherent 2D videos, but remain challenging in the physical world due to lack of 4D spatiotemporal scale. Typically, existing 4D generative models directly embed macro scale…

Computer Vision and Pattern Recognition · Computer Science 2026-05-11 Haonan Wang , Hanyu Zhou , Tao Gu , Luxin Yan

Co-speech gesture generation is crucial for creating lifelike avatars and enhancing human-computer interactions by synchronizing gestures with speech. Despite recent advancements, existing methods struggle with accurately identifying the…

Computer Vision and Pattern Recognition · Computer Science 2025-09-23 Pinxin Liu , Pengfei Zhang , Hyeongwoo Kim , Pablo Garrido , Ari Shapiro , Kyle Olszewski

Despite the astonishing progress in generative AI, 4D dynamic object generation remains an open challenge. With limited high-quality training data and heavy computing requirements, the combination of hallucinating unseen geometry together…

Computer Vision and Pattern Recognition · Computer Science 2025-04-18 Lu Sang , Zehranaz Canfes , Dongliang Cao , Riccardo Marin , Florian Bernard , Daniel Cremers

In recent years, there has been rapid development in 3D generation models, opening up new possibilities for applications such as simulating the dynamic movements of 3D objects and customizing their behaviors. However, current 3D generative…

Computer Vision and Pattern Recognition · Computer Science 2024-06-12 Fangfu Liu , Hanyang Wang , Shunyu Yao , Shengjun Zhang , Jie Zhou , Yueqi Duan

Video diffusion models lack explicit geometric supervision during training, leading to inconsistency artifacts such as object deformation, spatial drift, and depth violations in generated videos. To address this limitation, we propose a…

Computer Vision and Pattern Recognition · Computer Science 2026-03-24 Tengjiao Yin , Jinglei Shi , Heng Guo , Xi Wang

Reconstructing dynamic 4D scenes is an important yet challenging task. While 3D foundation models like VGGT excel in static settings, they often struggle with dynamic sequences where motion causes significant geometric ambiguity. To address…

Computer Vision and Pattern Recognition · Computer Science 2026-04-13 Ying Zang , Yidong Han , Chaotao Ding , Yuanqi Hu , Deyi Ji , Qi Zhu , Xuanfu Li , Jin Ma , Lingyun Sun , Tianrun Chen , Lanyun Zhu

Physical principles are fundamental to realistic visual simulation, but remain a significant oversight in transformer-based video generation. This gap highlights a critical limitation in rendering rigid body motion, a core tenet of…

Computer Vision and Pattern Recognition · Computer Science 2026-01-19 Qiyuan Zhang , Biao Gong , Shuai Tan , Zheng Zhang , Yujun Shen , Xing Zhu , Yuyuan Li , Kelu Yao , Chunhua Shen , Changqing Zou

Recent advancements in video diffusion models have shown exceptional abilities in simulating real-world dynamics and maintaining 3D consistency. This progress inspires us to investigate the potential of these models to ensure dynamic…

Computer Vision and Pattern Recognition · Computer Science 2024-12-11 Jianhong Bai , Menghan Xia , Xintao Wang , Ziyang Yuan , Xiao Fu , Zuozhu Liu , Haoji Hu , Pengfei Wan , Di Zhang

Recent video generation models have achieved remarkable progress and are now deployed in film, social media production, and advertising. Beyond their creative potential, such models also hold promise as world simulators for robotics and…

Computer Vision and Pattern Recognition · Computer Science 2026-03-24 David Romero , Ariana Bermudez , Viacheslav Iablochnikov , Hao Li , Fabio Pizzati , Ivan Laptev

Accurately simulating existing 3D objects and a wide variety of materials often demands expert knowledge and time-consuming physical parameter tuning to achieve the desired dynamic behavior. We introduce MotionPhysics, an end-to-end…

Computer Vision and Pattern Recognition · Computer Science 2026-01-05 Miaowei Wang , Jakub Zadrożny , Oisin Mac Aodha , Amir Vaxman

Human video synthesis aims to create lifelike characters in various environments, with wide applications in VR, storytelling, and content creation. While 2D diffusion-based methods have made significant progress, they struggle to generalize…

Computer Vision and Pattern Recognition · Computer Science 2024-12-19 Liyuan Cui , Xiaogang Xu , Wenqi Dong , Zesong Yang , Hujun Bao , Zhaopeng Cui

Computer vision is largely based on 2D techniques, with 3D vision still relegated to a relatively narrow subset of applications. However, by building on recent advances in 3D models such as neural radiance fields, some authors have shown…

Computer Vision and Pattern Recognition · Computer Science 2025-06-24 Vadim Tschernezki , Diane Larlus , Iro Laina , Andrea Vedaldi

Video generation is a challenging task that requires modeling plausible spatial and temporal dynamics in a video. Inspired by how humans perceive a video by grouping a scene into moving and stationary components, we propose a method that…

Computer Vision and Pattern Recognition · Computer Science 2022-03-29 Arti Keshari , Sonam Gupta , Sukhendu Das

We consider the problem of forecasting motion from a single image, i.e., predicting how objects in the world are likely to move, without the ability to observe other parameters such as the object velocities or the forces applied to them. We…

Computer Vision and Pattern Recognition · Computer Science 2026-05-26 Gabrijel Boduljak , Laurynas Karazija , Iro Laina , Christian Rupprecht , Andrea Vedaldi

Video world models can generate realistic futures from a single instruction, but they often fail to preserve consistent point-level motion over time. As a result, the generated videos appear plausible, yet lack the physical grounding…

Computer Vision and Pattern Recognition · Computer Science 2026-05-25 Kaichen Zhou , Yuzhen Chen , Fangneng Zhan , Hang Hua , Grace Chen , Xinhai Chang , Ao Qu , Yilun Du , Zhuang Liu , Paul Pu Liang , Mengyu Wang

One compelling application of artificial intelligence is to generate a video of a target person performing arbitrary desired motion (from a source person). While the state-of-the-art methods are able to synthesize a video demonstrating…

Computer Vision and Pattern Recognition · Computer Science 2022-12-26 Zhenguang Liu , Sifan Wu , Chejian Xu , Xiang Wang , Lei Zhu , Shuang Wu , Fuli Feng