English
Related papers

Related papers: GEM-4D: Geometry-Enhanced Video World Models for R…

200 papers

At its core, robotic manipulation is a problem of vision-to-geometry mapping ($f(v) \rightarrow G$). Physical actions are fundamentally defined by geometric properties like 3D positions and spatial relationships. Consequently, we argue that…

Robotics · Computer Science 2026-04-15 Zijian Song , Qichang Li , Jiawei Zhou , Zhenlong Yuan , Tianshui Chen , Liang Lin , Guangrun Wang

Recent advancements in human video synthesis have enabled the generation of high-quality videos through the application of stable diffusion models. However, existing methods predominantly concentrate on animating solely the human element…

Computer Vision and Pattern Recognition · Computer Science 2024-05-29 Jinlin Liu , Kai Yu , Mengyang Feng , Xiefan Guo , Miaomiao Cui

Video generation models have advanced rapidly and are beginning to show a strong understanding of physical dynamics. In this paper, we investigate how far an advanced video generation model such as Veo-3 can support generalizable robotic…

Robotics · Computer Science 2026-04-07 Zhongru Zhang , Chenghan Yang , Qingzhou Lu , Yanjiang Guo , Jianke Zhang , Yucheng Hu , Jianyu Chen

While Multimodal Large Language Models demonstrate impressive semantic capabilities, they often suffer from spatial blindness, struggling with fine-grained geometric reasoning and physical dynamics. Existing solutions typically rely on…

Computer Vision and Pattern Recognition · Computer Science 2026-03-20 Xianjin Wu , Dingkang Liang , Tianrui Feng , Kui Xia , Yumeng Zhang , Xiaofan Li , Xiao Tan , Xiang Bai

Several recent works have directly extended the image masked autoencoder (MAE) with random masking into video domain, achieving promising results. However, unlike images, both spatial and temporal information are important for video…

Computer Vision and Pattern Recognition · Computer Science 2023-08-25 David Fan , Jue Wang , Shuai Liao , Yi Zhu , Vimal Bhat , Hector Santos-Villalobos , Rohith MV , Xinyu Li

We introduce GeCo, a geometry-grounded metric for jointly detecting geometric deformation and occlusion-inconsistency artifacts in static scenes. By fusing residual motion and depth priors, GeCo produces interpretable, dense consistency…

Computer Vision and Pattern Recognition · Computer Science 2026-04-24 Leslie Gu , Junhwa Hur , Charles Herrmann , Fangneng Zhan , Todd Zickler , Deqing Sun , Hanspeter Pfister

Vision-Language-Action (VLA) models have emerged as a promising approach for enabling robots to follow language instructions and predict corresponding actions. However, current VLA models mainly rely on 2D visual inputs, neglecting the rich…

Robotics · Computer Science 2025-08-14 Lin Sun , Bin Xie , Yingfei Liu , Hao Shi , Tiancai Wang , Jiale Cao

Monocular 3D shape recovery is fundamental to geometric understanding, yet achieving robust generalization across arbitrary viewpoints and unseen object categories remains a significant challenge. In this paper, we present a generalizable…

Computer Vision and Pattern Recognition · Computer Science 2026-05-29 Yiyao Ma , Kai Chen , Zhongxiang Zhou , Zhuheng Song , Dongsheng Xie , Zelong Tan , Rong Xiong , Qi Dou

The integration of geometric reconstruction and generative modeling remains a critical challenge in developing AI systems capable of human-like spatial reasoning. This paper proposes Aether, a unified framework that enables geometry-aware…

Computer Vision and Pattern Recognition · Computer Science 2025-07-29 Aether Team , Haoyi Zhu , Yifan Wang , Jianjun Zhou , Wenzheng Chang , Yang Zhou , Zizun Li , Junyi Chen , Chunhua Shen , Jiangmiao Pang , Tong He

Action-conditioned video prediction models (often referred to as world models) have shown strong potential for robotics applications, but existing approaches are often slow and struggle to capture physically consistent interactions over…

Effective robotic manipulation relies on a precise understanding of 3D scene geometry, and one of the most straightforward ways to acquire such geometry is through multi-view observations. Motivated by this, we present GP3 -- a 3D…

Robotics · Computer Science 2025-09-22 Quanhao Qian , Guoyang Zhao , Gongjie Zhang , Jiuniu Wang , Ran Xu , Junlong Gao , Deli Zhao

Embodied action planning is a core challenge in robotics, requiring models to generate precise actions from visual observations and language instructions. While video generation world models are promising, their reliance on pixel-level…

Computer Vision and Pattern Recognition · Computer Science 2026-02-11 Yangcheng Yu , Xin Jin , Yu Shang , Xin Zhang , Haisheng Su , Wei Wu , Yong Li

Recent progress of video diffusion models have enabled extensive simulation of the physical world. While simulation with hand object interaction has been less explored. We propose DexSIM, a dexterous simulation framework for simulating…

Computer Vision and Pattern Recognition · Computer Science 2026-05-26 Adam Lee

Learning robust and scalable visual representations from massive multi-view video data remains a challenge in computer vision and autonomous driving. Existing pre-training methods either rely on expensive supervised learning with 3D…

Computer Vision and Pattern Recognition · Computer Science 2024-03-14 Jialv Zou , Bencheng Liao , Qian Zhang , Wenyu Liu , Xinggang Wang

Constructing 4D language fields is crucial for embodied AI, augmented/virtual reality, and 4D scene understanding, as they provide enriched semantic representations of dynamic environments and enable open-vocabulary querying in complex…

Computer Vision and Pattern Recognition · Computer Science 2025-12-05 Xianfeng Wu , Yajing Bai , Minghan Li , Xianzu Wu , Xueqi Zhao , Zhongyuan Lai , Wenyu Liu , Xinggang Wang

We propose DriveAnyMesh, a method for driving mesh guided by monocular video. Current 4D generation techniques encounter challenges with modern rendering engines. Implicit methods have low rendering efficiency and are unfriendly to…

Computer Vision and Pattern Recognition · Computer Science 2025-06-10 Yahao Shi , Yang Liu , Yanmin Wu , Xing Liu , Chen Zhao , Jie Luo , Bin Zhou

World models that support controllable and editable spatiotemporal environments are valuable for robotics, enabling scalable training data, repro ducible evaluation, and flexible task design. While recent text-to-video models generate…

Computer Vision and Pattern Recognition · Computer Science 2025-10-07 Xuehai He , Shijie Zhou , Thivyanth Venkateswaran , Kaizhi Zheng , Ziyu Wan , Achuta Kadambi , Xin Eric Wang

World modeling has become a cornerstone in AI research, enabling agents to understand, represent, and predict the dynamic environments they inhabit. While prior work largely emphasizes generative methods for 2D image and video data, they…

Video generative models pre-trained on large-scale internet datasets have achieved remarkable success, excelling at producing realistic synthetic videos. However, they often generate clips based on static prompts (e.g., text or images),…

Computer Vision and Pattern Recognition · Computer Science 2025-02-13 Haoran He , Yang Zhang , Liang Lin , Zhongwen Xu , Ling Pan

Video-guided 3D animation holds immense potential for content creation, offering intuitive and precise control over dynamic assets. However, practical deployment faces a critical yet frequently overlooked hurdle: the pose misalignment…

Computer Vision and Pattern Recognition · Computer Science 2026-05-15 Zijie Wu , Lixin Xu , Puhua Jiang , Sicong Liu , Chunchao Guo , Xiang Bai