English
Related papers

Related papers: X-World: Controllable Ego-Centric Multi-Camera Wor…

200 papers

Large-scale video generative models can synthesize diverse and realistic visual content for dynamic world creation, but they often lack element-wise controllability, hindering their use in editing scenes and training embodied AI agents. We…

Computer Vision and Pattern Recognition · Computer Science 2025-06-10 Sicheng Mo , Ziyang Leng , Leon Liu , Weizhen Wang , Honglin He , Bolei Zhou

Talking head generation is to generate video based on a given source identity and target motion. However, current methods face several challenges that limit the quality and controllability of the generated videos. First, the generated face…

Computer Vision and Pattern Recognition · Computer Science 2023-11-03 Yue Gao , Yuan Zhou , Jinglu Wang , Xiao Li , Xiang Ming , Yan Lu

The rise of multi-modal large language models(MLLMs) has spurred their applications in autonomous driving. Recent MLLM-based methods perform action by learning a direct mapping from perception to action, neglecting the dynamics of the world…

Computer Vision and Pattern Recognition · Computer Science 2024-09-06 Julong Wei , Shanshuai Yuan , Pengfei Li , Qingda Hu , Zhongxue Gan , Wenchao Ding

Cities, as the essential environment of human life, encompass diverse physical elements such as buildings, roads and vegetation, which continuously interact with dynamic entities like people and vehicles. Crafting realistic, interactive 3D…

Computer Vision and Pattern Recognition · Computer Science 2024-10-23 Yu Shang , Yuming Lin , Yu Zheng , Hangyu Fan , Jingtao Ding , Jie Feng , Jiansheng Chen , Li Tian , Yong Li

Egocentric vision is essential for both human and machine visual understanding, particularly in capturing the detailed hand-object interactions needed for manipulation tasks. Translating third-person views into first-person views…

Computer Vision and Pattern Recognition · Computer Science 2026-03-05 Junho Park , Andrew Sangwoo Ye , Taein Kwon

Building an efficient and physically consistent world model from limited observations is a long standing challenge in vision and robotics. Many existing world modeling pipelines are based on implicit generative models, which are hard to…

Computer Vision and Pattern Recognition · Computer Science 2025-06-06 Wenhao Hu , Xuexiang Wen , Xi Li , Gaoang Wang

Embodied agents in household environments must plan under partial observation: they need to remember objects, track state changes, and recover when actions fail. Existing benchmarks only partially test this ability. Egocentric video…

Artificial Intelligence · Computer Science 2026-05-14 Qinchuan Cheng , Zhantao Gong , Pengzhan Sun , Angela Yao , Xulei Yang , Shijie Li

Recent progress in 3D reconstruction has made it easy to create realistic digital twins from everyday environments. However, current digital twins remain largely static and are limited to navigation and view synthesis without embodied…

Computer Vision and Pattern Recognition · Computer Science 2025-12-22 Byungjun Kim , Taeksoo Kim , Junyoung Lee , Hanbyul Joo

Scalability in terms of object density in a scene is a primary challenge in unsupervised sequential object-oriented representation learning. Most of the previous models have been shown to work only on scenes with a few objects. In this…

Machine Learning · Computer Science 2020-03-06 Jindong Jiang , Sepehr Janghorbani , Gerard de Melo , Sungjin Ahn

World models - generative models that simulate environment dynamics conditioned on past observations and actions - are gaining prominence in planning, simulation, and embodied AI. However, evaluating their rollouts remains a fundamental…

Human videos offer a scalable way to train robot manipulation policies, but lack the action labels needed by standard imitation learning algorithms. Existing cross-embodiment approaches try to map human motion to robot actions, but often…

Current video models fail as world model as they lack fine-graiend control. General-purpose household robots require real-time fine motor control to handle delicate tasks and urgent situations. In this work, we introduce fine-grained…

Computer Vision and Pattern Recognition · Computer Science 2025-10-03 Yichen Li , Antonio Torralba

We propose the use of latent space generative world models to address the covariate shift problem in autonomous driving. A world model is a neural network capable of predicting an agent's next state given past states and actions. By…

The ability to simulate the effects of future actions on the world is a crucial ability of intelligent embodied agents, enabling agents to anticipate the effects of their actions and make plans accordingly. While a large body of existing…

Computer Vision and Pattern Recognition · Computer Science 2025-05-12 Siyuan Zhou , Yilun Du , Yuncong Yang , Lei Han , Peihao Chen , Dit-Yan Yeung , Chuang Gan

With increasing automation in passenger vehicles, the study of safe and smooth occupant-vehicle interaction and control transitions is key. In this study, we focus on the development of contextual, semantically meaningful representations of…

Robotics · Computer Science 2021-07-26 Akshay Rangesh , Nachiket Deo , Ross Greer , Pujitha Gunaratne , Mohan M. Trivedi

In recent years, data-driven techniques have greatly advanced autonomous driving systems, but the need for rare and diverse training data remains a challenge, requiring significant investment in equipment and labor. World models, which…

Computer Vision and Pattern Recognition · Computer Science 2025-03-21 Haiguang Wang , Daqi Liu , Hongwei Xie , Haisong Liu , Enhui Ma , Kaicheng Yu , Limin Wang , Bing Wang

Multi-camera vehicle tracking is one of the most complicated tasks in Computer Vision as it involves distinct tasks including Vehicle Detection, Tracking, and Re-identification. Despite the challenges, multi-camera vehicle tracking has…

Computer Vision and Pattern Recognition · Computer Science 2022-04-18 Pirazh Khorramshahi , Vineet Shenoy , Michael Pack , Rama Chellappa

We introduce HY-World 2.0, a multi-modal world model framework that advances our prior project HY-World 1.0. HY-World 2.0 accommodates diverse input modalities, including text prompts, single-view images, multi-view images, and videos, and…

Recent progress in advanced driver assistance systems and the race towards autonomous vehicles is mainly driven by two factors: (1) increasingly sophisticated algorithms that interpret the environment around the vehicle and react…

Computer Vision and Pattern Recognition · Computer Science 2017-04-04 Marius Cordts , Timo Rehfeld , Lukas Schneider , David Pfeiffer , Markus Enzweiler , Stefan Roth , Marc Pollefeys , Uwe Franke

Large-scale endeavors like and widespread community efforts such as Open-X-Embodiment have contributed to growing the scale of robot demonstration data. However, there is still an opportunity to improve the quality, quantity, and diversity…

Robotics · Computer Science 2024-08-30 Jiafei Duan , Wentao Yuan , Wilbert Pumacay , Yi Ru Wang , Kiana Ehsani , Dieter Fox , Ranjay Krishna