English
Related papers

Related papers: Kinema4D: Kinematic 4D World Modeling for Spatiote…

200 papers

Real-to-Sim-to-Real technique is gaining increasing interest for robotic manipulation, as it can generate scalable data in simulation while having narrower sim-to-real gap. However, previous methods mainly focused on environment-level…

Robotics · Computer Science 2026-01-27 Yiming Wang , Ruogu Zhang , Minyang Li , Hao Shi , Junbo Wang , Deyi Li , Jieji Ren , Wenhai Liu , Weiming Wang , Hao-Shu Fang

Understanding dynamic 4D scenes from an egocentric perspective-modeling changes in 3D spatial structure over time-is crucial for human-machine interaction, autonomous navigation, and embodied intelligence. While existing egocentric datasets…

Computer Vision and Pattern Recognition · Computer Science 2025-11-18 Junsheng Huang , Shengyu Hao , Bocheng Hu , Hongwei Wang , Gaoang Wang

Learned world models hold significant potential for robotic manipulation, as they can serve as simulator for real-world interactions. While extensive progress has been made in 2D video-based world models, these approaches often lack…

Robotics · Computer Science 2025-10-13 Chuanrui Zhang , Zhengxian Wu , Guanxing Lu , Yansong Tang , Ziwei Wang

Modeling the 3D world from sensor data for simulation is a scalable way of developing testing and validation environments for robotic learning problems such as autonomous driving. However, manually creating or re-creating real-world-like…

Computer Vision and Pattern Recognition · Computer Science 2023-08-29 Bokui Shen , Xinchen Yan , Charles R. Qi , Mahyar Najibi , Boyang Deng , Leonidas Guibas , Yin Zhou , Dragomir Anguelov

Recent research has been increasingly focusing on developing 3D world models that simulate complex real-world scenarios. World models have found broad applications across various domains, including embodied AI, autonomous driving,…

Artificial Intelligence · Computer Science 2025-09-09 Yinglin Duan , Zhengxia Zou , Tongwei Gu , Wei Jia , Zhan Zhao , Luyi Xu , Xinzhu Liu , Yenan Lin , Hao Jiang , Kang Chen , Shuang Qiu

Existing 4D synthesis methods primarily focus on object-level generation or dynamic scene synthesis with limited novel views, restricting their ability to generate multi-view consistent and immersive dynamic 4D scenes. To address these…

Computer Vision and Pattern Recognition · Computer Science 2025-06-25 Junwei Zhou , Xueting Li , Lu Qi , Ming-Hsuan Yang

World models for deformable objects should recover not only geometry and appearance, but also underlying physical dynamics, interaction grounding, and material behavior. Learning such a model from real videos is challenging because…

Computer Vision and Pattern Recognition · Computer Science 2026-05-21 Can Li , Zhoujian Li , Ren Li , Jie Gu , Lei Lei , Jingmin Chen , Lei Sun

Large-scale video generative models can synthesize diverse and realistic visual content for dynamic world creation, but they often lack element-wise controllability, hindering their use in editing scenes and training embodied AI agents. We…

Computer Vision and Pattern Recognition · Computer Science 2025-06-10 Sicheng Mo , Ziyang Leng , Leon Liu , Weizhen Wang , Honglin He , Bolei Zhou

We present OrbiSim, a novel robotic simulation paradigm that redefines world models as a fully differentiable physics engine for embodied intelligence. Unlike prior world models that focus on unconstrained imagination in latent or visual…

Robotics · Computer Science 2026-05-19 Jiajian Li , Jingyuan Huang , Junru Gong , Qi Wang , Xiaokang Yang , Yunbo Wang

Recent video diffusion models have achieved impressive capabilities as large-scale generative world models. However, these models often struggle with fine-grained physical consistency, exhibiting physically implausible dynamics over time.…

Computer Vision and Pattern Recognition · Computer Science 2026-03-09 Haoran Lu , Shang Wu , Jianshu Zhang , Maojiang Su , Guo Ye , Chenwei Xu , Lie Lu , Pranav Maneriker , Fan Du , Manling Li , Zhaoran Wang , Han Liu

In the intersection of computer vision and robotic perception, 4D reconstruction of dynamic scenes serve as the critical bridge connecting low-level geometric sensing with high-level semantic understanding. We present DINO\_4D, introducing…

Computer Vision and Pattern Recognition · Computer Science 2026-04-14 Yiru Yang , Zhuojie Wu , Quentin Marguet , Nishant Kumar Singh , Max Schulthess

Recent advancements in driving world models enable controllable generation of high-quality RGB videos or multimodal videos. Existing methods primarily focus on metrics related to generation quality and controllability. However, they often…

Computer Vision and Pattern Recognition · Computer Science 2026-03-10 Kai Zeng , Zhanqian Wu , Kaixin Xiong , Xiaobao Wei , Xiangyu Guo , Zhenxin Zhu , Kalok Ho , Lijun Zhou , Bohan Zeng , Ming Lu , Haiyang Sun , Bing Wang , Guang Chen , Hangjun Ye , Wentao Zhang

Being able to reproduce physical phenomena ranging from light interaction to contact mechanics, simulators are becoming increasingly useful in more and more application domains where real-world interaction or labeled data are difficult to…

Robotics · Computer Science 2022-09-13 Eric Heiden , Ziang Liu , Vibhav Vineet , Erwin Coumans , Gaurav S. Sukhatme

Simulation plays a key role in scaling robot learning and validating policies, but constructing simulations remains a labor-intensive process. This paper introduces ReGen, a generative simulation framework that automates simulation design…

Modeling dynamic 3D scenes is challenging due to their high-dimensional nature, which requires aggregating information from multiple views to reconstruct time-evolving 3D geometry and motion. We present a novel multi-video 4D Gaussian…

Computer Vision and Pattern Recognition · Computer Science 2025-12-05 Yonghan Lee , Tsung-Wei Huang , Shiv Gehlot , Jaehoon Choi , Guan-Ming Su , Dinesh Manocha

Recent video generators achieve striking photorealism, yet remain fundamentally inconsistent in 3D. We present WorldReel, a 4D video generator that is natively spatio-temporally consistent. WorldReel jointly produces RGB frames together…

Computer Vision and Pattern Recognition · Computer Science 2025-12-09 Shaoheng Fang , Hanwen Jiang , Yunpeng Bai , Niloy J. Mitra , Qixing Huang

World models for interactive video generation have largely focused on single-agent settings, where future observations are generated from a single control signal. However, many generated environments require multi-agent interaction:…

Computer Vision and Pattern Recognition · Computer Science 2026-05-28 Fangfu Liu , Kai He , Tianchang Shen , Tianshi Cao , Sanja Fidler , Yueqi Duan , Jun Gao , Igor Gilitschenski , Zian Wang , Xuanchi Ren

Robot manipulation in the real world is fundamentally constrained by the visual sim2real gap, where depth observations collected in simulation fail to reflect the complex noise patterns inherent to real sensors. In this work, inspired by…

Robotics · Computer Science 2025-12-09 Xiujian Liang , Jiacheng Liu , Mingyang Sun , Qichen He , Cewu Lu , Jianhua Sun

Humans develop an understanding of intuitive physics through active interaction with the world. This approach is in stark contrast to current video models, such as Sora, which rely on passive observation and therefore struggle with grasping…

Extended reality (XR) demands generative models that respond to users' tracked real-world motion, yet current video world models accept only coarse control signals such as text or keyboard input, limiting their utility for embodied…

Computer Vision and Pattern Recognition · Computer Science 2026-02-23 Linxi Xie , Lisong C. Sun , Ashley Neall , Tong Wu , Shengqu Cai , Gordon Wetzstein
‹ Prev 1 4 5 6 7 8 10 Next ›