English
Related papers

Related papers: Kinema4D: Kinematic 4D World Modeling for Spatiote…

200 papers

Directly learning to model 4D content, including shape, color, and motion, is challenging. Existing methods rely on pose priors for motion control, resulting in limited motion diversity and continuity in details. To address this, we propose…

Computer Vision and Pattern Recognition · Computer Science 2025-06-17 Qitong Yang , Mingtao Feng , Zijie Wu , Shijie Sun , Weisheng Dong , Yaonan Wang , Ajmal Mian

Realizing scaling laws in embodied AI has become a focus. However, previous work has been scattered across diverse simulation platforms, with assets and models lacking unified interfaces, which has led to inefficiencies in research. To…

3D scene generation has garnered growing attention in recent years and has made significant progress. Generating 4D cities is more challenging than 3D scenes due to the presence of structurally complex, visually diverse objects like…

Computer Vision and Pattern Recognition · Computer Science 2025-09-03 Haozhe Xie , Zhaoxi Chen , Fangzhou Hong , Ziwei Liu

Generative artificial intelligence has recently progressed from static image and video synthesis to 3D content generation, culminating in the emergence of 4D generation-the task of synthesizing temporally coherent dynamic 3D assets guided…

Computer Vision and Pattern Recognition · Computer Science 2025-07-25 Qiaowei Miao , Kehan Li , Jinsheng Quan , Zhiyuan Min , Shaojie Ma , Yichao Xu , Yi Yang , Ping Liu , Yawei Luo

Training robot policies within a learned world model is trending due to the inefficiency of real-world interactions. The established image-based world models and policies have shown prior success, but lack robust geometric information that…

Robotics · Computer Science 2025-09-18 Guanxing Lu , Baoxiong Jia , Puhao Li , Yixin Chen , Ziwei Wang , Yansong Tang , Siyuan Huang

Physical embodiment is a required component for robots that are structurally coupled with their real-world environments. However, most socially interactive robots do not need to physically interact with their environments in order to…

Robotics · Computer Science 2019-12-03 Eric Deng , Bilge Mutlu , Maja Mataric

The ability to automatically generate large-scale, interactive, and physically realistic 3D environments is crucial for advancing robotic learning and embodied intelligence. However, existing generative approaches often fail to capture the…

Computer Vision and Pattern Recognition · Computer Science 2026-01-19 ChunTeng Chen , YiChen Hsu , YiWen Liu , WeiFang Sun , TsaiChing Ni , ChunYi Lee , Min Sun , YuanFu Yang

Traditional approaches to quadruped control frequently employ simplified, hand-derived models. This significantly reduces the capability of the robot since its effective kinematic range is curtailed. In addition, kinodynamic constraints are…

Modern robotic manipulation primarily relies on visual observations in a 2D color space for skill learning but suffers from poor generalization. In contrast, humans, living in a 3D world, depend more on physical properties-such as distance,…

4D driving simulation is essential for developing realistic autonomous driving simulators. Despite advancements in existing methods for generating driving scenes, significant challenges remain in view transformation and spatial-temporal…

Computer Vision and Pattern Recognition · Computer Science 2024-12-12 Lening Wang , Wenzhao Zheng , Dalong Du , Yunpeng Zhang , Yilong Ren , Han Jiang , Zhiyong Cui , Haiyang Yu , Jie Zhou , Jiwen Lu , Shanghang Zhang

Embodied AI research is undergoing a shift toward vision-centric perceptual paradigms. While massively parallel simulators have catalyzed breakthroughs in proprioception-based locomotion, their potential remains largely untapped for…

World models allow autonomous agents to plan and explore by predicting the visual outcomes of different actions. However, for robot manipulation, it is challenging to accurately model the fine-grained robot-object interaction within the…

Robotics · Computer Science 2025-07-30 Fangqi Zhu , Hongtao Wu , Song Guo , Yuxiao Liu , Chilam Cheang , Tao Kong

Data-driven approaches have revolutionized 3D vision, enabling transformers to effectively reconstruct and generate static 3D objects. However, generating simulative 4D dynamics -- realistic temporal deformations of static objects under…

Computer Vision and Pattern Recognition · Computer Science 2026-05-29 Chen Geng , Guangzhao He , Yue Gao , Yunzhi Zhang , Shangzhe Wu , Jiajun Wu

Video-based world models offer a powerful paradigm for embodied simulation and planning, yet state-of-the-art models often generate physically implausible manipulations - such as object penetration and anti-gravity motion - due to training…

Computer Vision and Pattern Recognition · Computer Science 2026-03-30 Yuzhi Chen , Ronghan Chen , Dongjie Huo , Yandan Yang , Dekang Qi , Haoyun Liu , Tong Lin , Shuang Zeng , Junjin Xiao , Xinyuan Chang , Feng Xiong , Xing Wei , Zhiheng Ma , Mu Xu

We present Free4D, a novel tuning-free framework for 4D scene generation from a single image. Existing methods either focus on object-level generation, making scene-level generation infeasible, or rely on large-scale multi-view video…

Computer Vision and Pattern Recognition · Computer Science 2025-03-27 Tianqi Liu , Zihao Huang , Zhaoxi Chen , Guangcong Wang , Shoukang Hu , Liao Shen , Huiqiang Sun , Zhiguo Cao , Wei Li , Ziwei Liu

Human-scene interaction (HSI) generation is crucial for applications in embodied AI, virtual reality, and robotics. Yet, existing methods cannot synthesize interactions in unseen environments such as in-the-wild scenes or reconstructed…

Computer Vision and Pattern Recognition · Computer Science 2025-03-24 Hongjie Li , Hong-Xing Yu , Jiaman Li , Jiajun Wu

Robotic manipulation with deformable objects represents a data-intensive regime in embodied learning, where shape, contact, and topology co-evolve in ways that far exceed the variability of rigids. Although simulation promises relief from…

We introduce LivingWorld, an interactive framework for generating 4D worlds with environmental dynamics from a single image. While recent advances in 3D scene generation enable large-scale environment creation, most approaches focus…

Computer Vision and Pattern Recognition · Computer Science 2026-04-03 Hyeongju Mun , In-Hwan Jin , Sohyeong Kim , Kyeongbo Kong

We introduce EgoSim, a closed-loop egocentric world simulator that generates spatially consistent interaction videos and persistently updates the underlying 3D scene state for continuous simulation. Existing egocentric simulators either…

Computer Vision and Pattern Recognition · Computer Science 2026-04-02 Jinkun Hao , Mingda Jia , Ruiyan Wang , Xihui Liu , Ran Yi , Lizhuang Ma , Jiangmiao Pang , Xudong Xu

The landscape of video generation is shifting, from a focus on generating visually appealing clips to building virtual environments that support interaction and maintain physical plausibility. These developments point toward the emergence…

Artificial Intelligence · Computer Science 2026-02-09 Jingtong Yue , Ziqi Huang , Zhaoxi Chen , Xintao Wang , Pengfei Wan , Ziwei Liu