中文
相关论文

相关论文: MVISTA-4D: View-Consistent 4D World Model with Tes…

200 篇论文

World models aim to endow AI systems with the ability to represent, generate, and interact with dynamic environments in a coherent and temporally consistent manner. While recent video generation models have demonstrated impressive visual…

World models have recently re-emerged as a central paradigm for embodied intelligence, robotics, autonomous driving, and model-based reinforcement learning. However, current world model research is often dominated by three partially…

人工智能 · 计算机科学 2026-05-27 Sen Cui , Jingheng Ma

Reliable anticipation of traffic accidents is essential for advancing autonomous driving systems. However, this objective is limited by two fundamental challenges: the scarcity of diverse, high-quality training data and the frequent absence…

计算机视觉与模式识别 · 计算机科学 2025-07-18 Yanchen Guan , Haicheng Liao , Chengyue Wang , Xingcheng Liu , Jiaxun Zhang , Zhenning Li

Generating multi-view videos for autonomous driving training has recently gained much attention, with the challenge of addressing both cross-view and cross-frame consistency. Existing methods typically apply decoupled attention mechanisms…

计算机视觉与模式识别 · 计算机科学 2024-12-10 Hannan Lu , Xiaohe Wu , Shudong Wang , Xiameng Qin , Xinyu Zhang , Junyu Han , Wangmeng Zuo , Ji Tao

We propose X-WAM, a Unified 4D World Model that unifies real-time robotic action execution and high-fidelity 4D world synthesis (video + 3D reconstruction) in a single framework, addressing the critical limitations of prior unified world…

机器人学 · 计算机科学 2026-05-08 Jun Guo , Qiwei Li , Peiyan Li , Zilong Chen , Nan Sun , Yifei Su , Heyun Wang , Yuan Zhang , Xinghang Li , Huaping Liu

Despite increasingly realistic image quality, recent 3D image generative models often operate on 3D volumes of fixed extent with limited camera motions. We investigate the task of unconditionally synthesizing unbounded nature scenes,…

计算机视觉与模式识别 · 计算机科学 2023-03-24 Lucy Chai , Richard Tucker , Zhengqi Li , Phillip Isola , Noah Snavely

Video generative models have emerged as a promising robotics backbone, capable of generating videos that depict the completion of complex tasks across embodiments and environments. Recent work proposes robot foundation models that jointly…

机器人学 · 计算机科学 2026-05-28 Sizhe Lester Li , Evan Kim , Xingjian Bai , Tong Zhao , Tao Pang , Max Simchowitz , Vincent Sitzmann

Pretrained video diffusion models provide powerful spatiotemporal generative priors, making them a natural foundation for robotic world models. While recent world-action models jointly optimize future videos and actions, they predominantly…

计算机视觉与模式识别 · 计算机科学 2026-05-08 Zhaoyang Yang , Yurun Jin , Lizhe Qi , Cong Huang , Kai Chen

Building an efficient and physically consistent world model from limited observations is a long standing challenge in vision and robotics. Many existing world modeling pipelines are based on implicit generative models, which are hard to…

计算机视觉与模式识别 · 计算机科学 2025-06-06 Wenhao Hu , Xuexiang Wen , Xi Li , Gaoang Wang

Recent advancements in foundation models for 2D vision have substantially improved the analysis of dynamic scenes from monocular videos. However, despite their strong generalization capabilities, these models often lack 3D consistency, a…

计算机视觉与模式识别 · 计算机科学 2025-12-04 Haoran Zhou , Gim Hee Lee

How to build AI that understands human intentions, and uses this knowledge to collaborate with people? We describe a computational framework for evaluating models of goal inference in the domain of 3D motor actions, which receives as input…

人工智能 · 计算机科学 2021-12-03 Yingdong Qian , Marta Kryven , Tao Gao , Hanbyul Joo , Josh Tenenbaum

Aerodynamic inverse design can improve vehicle and aircraft efficiency, but practical design rarely seeks performance alone: vehicle refinement must reduce drag while preserving visual features linked to design language, brand recognition…

机器学习 · 计算机科学 2026-05-29 Huaguan Chen , Ning Lin , Luxi Chen , Jiacheng Cen , Rui Zhang , Wenbing Huang , Chongxuan Li , Hao Sun

Deformable objects manipulation can benefit from representations that seamlessly integrate vision and touch while handling occlusions. In this work, we present a novel approach for, and real-world demonstration of, multimodal visuo-tactile…

机器人学 · 计算机科学 2022-10-10 Youngsun Wi , Andy Zeng , Pete Florence , Nima Fazeli

Relational object rearrangement (ROR) tasks (e.g., insert flower to vase) require a robot to manipulate objects with precise semantic and geometric reasoning. Existing approaches either rely on pre-collected demonstrations that struggle to…

机器人学 · 计算机科学 2025-09-23 Liang Heng , Jiadong Xu , Yiwen Wang , Xiaoqi Li , Muhe Cai , Yan Shen , Juan Zhu , Guanghui Ren , Hao Dong

We introduce a pipeline that enhances a general-purpose Vision Language Model, GPT-4V(ision), to facilitate one-shot visual teaching for robotic manipulation. This system analyzes videos of humans performing tasks and outputs executable…

机器人学 · 计算机科学 2024-10-11 Naoki Wake , Atsushi Kanehira , Kazuhiro Sasabuchi , Jun Takamatsu , Katsushi Ikeuchi

Video world models aim to simulate dynamic, real-world environments, yet existing methods struggle to provide unified and precise control over camera and multi-object motion, as videos inherently capture dynamics in the projected 2D image…

计算机视觉与模式识别 · 计算机科学 2026-03-31 Sixiao Zheng , Minghao Yin , Wenbo Hu , Xiaoyu Li , Ying Shan , Yanwei Fu

In the paper, we propose a robust real-time visual odometry in dynamic environments via rigid-motion model updated by scene flow. The proposed algorithm consists of spatial motion segmentation and temporal motion tracking. The spatial…

机器人学 · 计算机科学 2019-07-22 Sangil Lee , Clark Youngdong Son , H. Jin Kim

A key challenge for an agent learning to interact with the world is to reason about physical properties of objects and to foresee their dynamics under the effect of applied forces. In order to scale learning through interaction to many…

机器人学 · 计算机科学 2020-08-04 Iman Nematollahi , Oier Mees , Lukas Hermann , Wolfram Burgard

World models predict future transitions from observations and actions. Existing works predominantly focus on image generation only. Visual feature-based world models, on the other hand, predict future visual features instead of raw video…

计算机视觉与模式识别 · 计算机科学 2026-05-11 Xinyu Zhang , Zhengtong Xu , Yutian Tao , Yeping Wang , Yu She , Abdeslam Boularias

A world model is an internal model that simulates how the world evolves. Given past observations and actions, it predicts the future physical state of both the embodied agent and its environment. Accurate world models are essential for…

机器学习 · 计算机科学 2026-04-22 Zaishuo Xia , Yukuan Lu , Xinyi Li , Yifan Xu , Yubei Chen