English
Related papers

Related papers: AirScape: An Aerial Generative World Model with Mo…

200 papers

We introduce EnerVerse, a generative robotics foundation model that constructs and interprets embodied spaces. EnerVerse employs a chunk-wise autoregressive video diffusion framework to predict future embodied spaces from instructions,…

In the realm of computer vision and robotics, embodied agents are expected to explore their environment and carry out human instructions. This necessitates the ability to fully understand 3D scenes given their first-person observations and…

Computer Vision and Pattern Recognition · Computer Science 2023-12-27 Tai Wang , Xiaohan Mao , Chenming Zhu , Runsen Xu , Ruiyuan Lyu , Peisen Li , Xiao Chen , Wenwei Zhang , Kai Chen , Tianfan Xue , Xihui Liu , Cewu Lu , Dahua Lin , Jiangmiao Pang

Enabling embodied agents to imagine future states is essential for robust and generalizable visual navigation. Yet, state-of-the-art systems typically rely on modular designs that decouple navigation planning from visual world modeling,…

Artificial Intelligence · Computer Science 2026-03-24 Yifei Dong , Fengyi Wu , Guangyu Chen , Lingdong Kong , Xu Zhu , Qiyu Hu , Yuxuan Zhou , Jingdong Sun , Jun-Yan He , Qi Dai , Alexander G. Hauptmann , Zhi-Qi Cheng

The ability to decompose complex multi-object scenes into meaningful abstractions like objects is fundamental to achieve higher-level cognition. Previous approaches for unsupervised object-oriented scene representation learning are either…

Machine Learning · Computer Science 2020-03-17 Zhixuan Lin , Yi-Fu Wu , Skand Vishwanath Peri , Weihao Sun , Gautam Singh , Fei Deng , Jindong Jiang , Sungjin Ahn

Generative world models (WMs) can now simulate worlds with striking visual realism, which naturally raises the question of whether they can endow embodied agents with predictive perception for decision making. Progress on this question has…

When exploring new areas, robotic systems generally exclusively plan and execute controls over geometry that has been directly measured. When entering space that was previously obstructed from view such as turning corners in hallways or…

Robotics · Computer Science 2024-03-19 Alec Reed , Brendan Crowe , Doncey Albin , Lorin Achey , Bradley Hayes , Christoffer Heckman

Predicting future human pose is a fundamental application for machine intelligence, which drives robots to plan their behavior and paths ahead of time to seamlessly accomplish human-robot collaboration in real-world 3D scenarios. Despite…

Computer Vision and Pattern Recognition · Computer Science 2024-05-07 Zhenyu Lou , Qiongjie Cui , Haofan Wang , Xu Tang , Hong Zhou

Motion prediction in unstructured environments is a difficult problem and is essential for safe and efficient human-robot space sharing and collaboration. In this work, we focus on manipulation movements in environments such as homes,…

Robotics · Computer Science 2020-07-21 Philipp Kratzer , Niteesh Balachandra Midlagajni , Marc Toussaint , Jim Mainprice

In computer vision, video-based approaches have been widely explored for the early classification and the prediction of actions or activities. However, it remains unclear whether this modality (as compared to 3D kinematics) can still be…

Computer Vision and Pattern Recognition · Computer Science 2017-08-04 Andrea Zunino , Jacopo Cavazza , Atesh Koul , Andrea Cavallo , Cristina Becchio , Vittorio Murino

We release two artificial datasets, Simulated Flying Shapes and Simulated Planar Manipulator that allow to test the learning ability of video processing systems. In particular, the dataset is meant as a tool which allows to easily assess…

Computer Vision and Pattern Recognition · Computer Science 2018-07-03 Fabio Ferreira , Jonas Rothfuss , Eren Erdal Aksoy , You Zhou , Tamim Asfour

A world model enables an intelligent agent to imagine, predict, and reason about how the world evolves in response to its actions, and accordingly to plan and strategize. While recent video generation models produce realistic visual…

Traditional control and planning for robotic manipulation heavily rely on precise physical models and predefined action sequences. While effective in structured environments, such approaches often fail in real-world scenarios due to…

Robotics · Computer Science 2025-08-08 Jin Wang , Weijie Wang , Boyuan Deng , Heng Zhang , Rui Dai , Nikos Tsagarakis

Recent generative AI models have achieved remarkable breakthroughs in language and visual understanding. However, although these models can generate realistic visual content, their spatial scale remains confined to bounded environments,…

Computer Vision and Pattern Recognition · Computer Science 2026-04-28 Jinqi Cao , Zhiping Yu , Baihong Lin , Chenyang Liu , Zhenwei Shi , Zhengxia Zou

Optimizing and refining action execution through exploration and interaction is a promising way for robotic manipulation. However, practical approaches to interaction-driven robotic learning are still underexplored, particularly for…

Robotics · Computer Science 2025-09-24 Yibo Peng , Jiahao Yang , Shenhao Yan , Ziyu Huang , Shuang Li , Shuguang Cui , Yiming Zhao , Yatong Han

Effective human-robot interaction requires robots to identify human intentions and generate expressive, socially appropriate motions in real-time. Existing approaches often rely on fixed motion libraries or computationally expensive…

Robotics · Computer Science 2025-09-30 Lingfan Bao , Yan Pan , Tianhu Peng , Dimitrios Kanoulas , Chengxu Zhou

We present FlightDiffusion, a diffusion-model-based framework for training autonomous drones from first-person view (FPV) video. Our model generates realistic video sequences from a single frame, enriched with corresponding action spaces to…

City-scale 3D generation is of great importance for the development of embodied intelligence and world models. Existing methods, however, face significant challenges regarding quality, fidelity, and scalability in 3D world generation. Thus,…

Computer Vision and Pattern Recognition · Computer Science 2025-11-25 Shengyuan Wang , Zhiheng Zheng , Yu Shang , Lixuan He , Yangcheng Yu , Fan Hangyu , Jie Feng , Qingmin Liao , Yong Li

Embodied reasoning is inherently viewpoint-dependent: what is visible, occluded, or reachable depends critically on where the agent stands. However, existing spatial memory systems for embodied agents typically store either multi-view…

Artificial Intelligence · Computer Science 2026-03-17 JooHyun Park , HyeongYeop Kang

Human behaviors in the real world naturally encode rich, long-term contextual information that can be leveraged to train embodied agents for perception, understanding, and acting. However, existing capture systems typically rely on costly…

Computer Vision and Pattern Recognition · Computer Science 2026-04-03 Wenjia Wang , Liang Pan , Huaijin Pi , Yuke Lou , Xuqian Ren , Yifan Wu , Zhouyingcheng Liao , Lei Yang , Rishabh Dabral , Christian Theobalt , Taku Komura

Shared control in teleoperation for providing robot assistance to accomplish object manipulation, called telemanipulation, is a new promising yet challenging problem. This has unique challenges--on top of teleoperation challenges in…

Robotics · Computer Science 2025-04-02 Michael Bowman , Jiucai Zhang , Xiaoli Zhang