English
Related papers

Related papers: GEM-4D: Geometry-Enhanced Video World Models for R…

200 papers

Identifying predictive world models for robots in novel environments from sparse online observations is essential for robot task planning and execution in novel environments. However, existing methods that leverage differentiable…

Robotics · Computer Science 2025-05-13 Yifan Zhu , Tianyi Xiang , Aaron Dollar , Zherong Pan

While the recent advances in Multimodal Large Language Models (MLLMs) constitute a significant leap forward in the field, these models are predominantly confined to the realm of input-side multimodal comprehension, lacking the capacity for…

Computer Vision and Pattern Recognition · Computer Science 2024-10-29 Zhanyu Wang , Longyue Wang , Zhen Zhao , Minghao Wu , Chenyang Lyu , Huayang Li , Deng Cai , Luping Zhou , Shuming Shi , Zhaopeng Tu

Recent video generation models have achieved remarkable progress and are now deployed in film, social media production, and advertising. Beyond their creative potential, such models also hold promise as world simulators for robotics and…

Computer Vision and Pattern Recognition · Computer Science 2026-03-24 David Romero , Ariana Bermudez , Viacheslav Iablochnikov , Hao Li , Fabio Pizzati , Ivan Laptev

In recent years, 2D Vision-Language Models (VLMs) have made significant strides in image-text understanding tasks. However, their performance in 3D spatial comprehension, which is critical for embodied intelligence, remains limited. Recent…

Computer Vision and Pattern Recognition · Computer Science 2025-03-12 Zhangyang Qi , Zhixiong Zhang , Ye Fang , Jiaqi Wang , Hengshuang Zhao

Recent developments in generative models and large-scale datasets have substantially advanced 3D world generation, facilitating a broad range of domains including spatial intelligence, embodied intelligence, and autonomous driving. While…

Computer Vision and Pattern Recognition · Computer Science 2026-05-14 Hanxin Zhu , Cong Wang , Peiyan Tu , Jiayi Luo , Tianyu He , Xin Jin , Zhibo Chen

We introduce PhysWorld, a framework that enables robot learning from video generation through physical world modeling. Recent video generation models can synthesize photorealistic visual demonstrations from language commands and images,…

We present a survey on 4D generation and reconstruction, a fast-evolving subfield of computer graphics whose developments have been propelled by recent advances in neural fields, geometric and motion deep learning, as well as 3D generative…

Computer Vision and Pattern Recognition · Computer Science 2026-03-10 Mingrui Zhao , Sauradip Nag , Kai Wang , Aditya Vora , Guangda Ji , Peter Chun , Ali Mahdavi-Amiri , Hao Zhang

Human Mesh Recovery (HMR) aims to reconstruct 3D human pose and shape from 2D observations and is fundamental to human-centric understanding in real-world scenarios. While recent image-based HMR methods such as SAM 3D Body achieve strong…

Computer Vision and Pattern Recognition · Computer Science 2025-12-10 Mingqi Gao , Yunqi Miao , Jungong Han

We present a novel algorithm for estimating the broad 3D geometric structure of outdoor video scenes. Leveraging spatio-temporal video segmentation, we decompose a dynamic scene captured by a video into geometric classes, based on…

Computer Vision and Pattern Recognition · Computer Science 2016-11-17 S. Hussain Raza , Matthias Grundmann , Irfan Essa

Bimanual manipulation requires policies that can reason about 3D geometry, anticipate how it evolves under action, and generate smooth, coordinated motions. However, existing methods typically rely on 2D features with limited spatial…

Computer Vision and Pattern Recognition · Computer Science 2026-03-02 Chongyang Xu , Haipeng Li , Shen Cheng , Jingyu Hu , Haoqiang Fan , Ziliang Feng , Shuaicheng Liu

Imitation learning has emerged as a promising approach towards building generalist robots. However, scaling imitation learning for large robot foundation models remains challenging due to its reliance on high-quality expert demonstrations.…

Robotics · Computer Science 2025-05-26 Chuning Zhu , Raymond Yu , Siyuan Feng , Benjamin Burchfiel , Paarth Shah , Abhishek Gupta

Dexterous manipulation remains a challenging robotics problem, largely due to the difficulty of collecting extensive human demonstrations for learning. In this paper, we introduce \textsc{Gen2Real}, which replaces costly human demos with…

Robotics · Computer Science 2025-09-18 Kai Ye , Yuhang Wu , Shuyuan Hu , Junliang Li , Meng Liu , Yongquan Chen , Rui Huang

The ability to simulate the effects of future actions on the world is a crucial ability of intelligent embodied agents, enabling agents to anticipate the effects of their actions and make plans accordingly. While a large body of existing…

Computer Vision and Pattern Recognition · Computer Science 2025-05-12 Siyuan Zhou , Yilun Du , Yuncong Yang , Lei Han , Peihao Chen , Dit-Yan Yeung , Chuang Gan

Video generation is experiencing rapid growth, driven by advances in diffusion models and the development of better and larger datasets. However, producing high-quality videos remains challenging due to the high-dimensional data and the…

Computer Vision and Pattern Recognition · Computer Science 2025-04-10 Elia Peruzzo , Dejia Xu , Xingqian Xu , Humphrey Shi , Nicu Sebe

We present Free4D, a novel tuning-free framework for 4D scene generation from a single image. Existing methods either focus on object-level generation, making scene-level generation infeasible, or rely on large-scale multi-view video…

Computer Vision and Pattern Recognition · Computer Science 2025-03-27 Tianqi Liu , Zihao Huang , Zhaoxi Chen , Guangcong Wang , Shoukang Hu , Liao Shen , Huiqiang Sun , Zhiguo Cao , Wei Li , Ziwei Liu

Closed-loop simulation is essential for advancing end-to-end autonomous driving systems. Contemporary sensor simulation methods, such as NeRF and 3DGS, rely predominantly on conditions closely aligned with training data distributions, which…

Computer Vision and Pattern Recognition · Computer Science 2024-11-26 Guosheng Zhao , Chaojun Ni , Xiaofeng Wang , Zheng Zhu , Xueyang Zhang , Yida Wang , Guan Huang , Xinze Chen , Boyuan Wang , Youyi Zhang , Wenjun Mei , Xingang Wang

A current limitation of video generative video models is that they generate plausible looking frames, but poor motion -- an issue that is not well captured by FVD and other popular methods for evaluating generated videos. Here we go beyond…

Previous text-to-4D methods have leveraged multiple Score Distillation Sampling (SDS) techniques, combining motion priors from video-based diffusion models (DMs) with geometric priors from multiview DMs to implicitly guide 4D renderings.…

Computer Vision and Pattern Recognition · Computer Science 2024-11-20 Qiaowei Miao , JinSheng Quan , Kehan Li , Yawei Luo

We present a novel video generation framework that integrates 3-dimensional geometry and dynamic awareness. To achieve this, we augment 2D videos with 3D point trajectories and align them in pixel space. The resulting 3D-aware video…

Computer Vision and Pattern Recognition · Computer Science 2025-10-24 Yunuo Chen , Junli Cao , Vidit Goel , Sergei Korolev , Chenfanfu Jiang , Jian Ren , Sergey Tulyakov , Anil Kag

We propose Mesh4D, a feed-forward model for monocular 4D mesh reconstruction. Given a monocular video of a dynamic object, our model reconstructs the object's complete 3D shape and motion, represented as a deformation field. Our key…

Computer Vision and Pattern Recognition · Computer Science 2026-01-09 Zeren Jiang , Chuanxia Zheng , Iro Laina , Diane Larlus , Andrea Vedaldi