English
Related papers

Related papers: iMoWM: Taming Interactive Multi-Modal World Model …

200 papers

Loco-manipulation is a fundamental challenge for humanoid robots to achieve versatile interactions in human environments. Although recent studies have made significant progress in humanoid whole-body control, loco-manipulation remains…

Robotics · Computer Science 2025-10-14 Yuhui Fu , Feiyang Xie , Chaoyi Xu , Jing Xiong , Haoqi Yuan , Zongqing Lu

World models for autonomous driving have the potential to dramatically improve the reasoning capabilities of today's systems. However, most works focus on camera data, with only a few that leverage lidar data or combine both to better…

Machine Learning · Computer Science 2025-08-21 Daniel Bogdoll , Yitian Yang , Tim Joseph , Melih Yazgan , J. Marius Zöllner

Integrating AI into the physical layer is a cornerstone of 6G networks. However, current data-driven approaches struggle to generalize across dynamic environments because they lack an intrinsic understanding of electromagnetic wave…

Networking and Internet Architecture · Computer Science 2026-03-27 Ziqi Chen , Yi Ren , Yixuan Huang , Qi Sun , Nan Li , Yuhong Huang , Chih-Lin I , Yifan Li , Liang Xia

Exploiting the promise of recent advances in imitation learning for mobile manipulation will require the collection of large numbers of human-guided demonstrations. This paper proposes an open-source design for an inexpensive, robust, and…

Robots in uncertain real-world environments must perform both goal-directed and exploratory actions. However, most deep learning-based control methods neglect exploration and struggle under uncertainty. To address this, we adopt deep active…

Robotics · Computer Science 2025-12-02 Kentaro Fujii , Shingo Murata

We present a deep imitation learning framework for robotic bimanual manipulation in a continuous state-action space. A core challenge is to generalize the manipulation skills to objects in different locations. We hypothesize that modeling…

Video world models have achieved remarkable success in simulating environmental dynamics in response to actions by users or agents. They are modeled as action-conditioned video generation models that take historical frames and current…

Computer Vision and Pattern Recognition · Computer Science 2026-04-22 Haoyu Wu , Jiwen Yu , Yingtian Zou , Xihui Liu

World modelling, i.e. building a representation of the rules that govern the world so as to predict its evolution, is an essential ability for any agent interacting with the physical world. Recent applications of the Transformer…

Machine Learning · Computer Science 2024-05-31 Francesco Petri , Luigi Asprino , Aldo Gangemi

Generative video models, a leading approach to world modeling, face fundamental limitations. They often violate physical and logical rules, lack interactivity, and operate as opaque black boxes ill-suited for building structured, queryable…

Computer Vision and Pattern Recognition · Computer Science 2025-12-15 Felix O'Mahony , Roberto Cipolla , Ayush Tewari

Generative world models (WMs) can now simulate worlds with striking visual realism, which naturally raises the question of whether they can endow embodied agents with predictive perception for decision making. Progress on this question has…

Multimodal representation learning has shown promising improvements on various vision-language tasks. Most existing methods excel at building global-level alignment between vision and language while lacking effective fine-grained image-text…

Computer Vision and Pattern Recognition · Computer Science 2023-06-16 Zijia Zhao , Longteng Guo , Xingjian He , Shuai Shao , Zehuan Yuan , Jing Liu

Navigation is a fundamental skill of agents with visual-motor capabilities. We introduce a Navigation World Model (NWM), a controllable video generation model that predicts future visual observations based on past observations and…

Computer Vision and Pattern Recognition · Computer Science 2025-04-15 Amir Bar , Gaoyue Zhou , Danny Tran , Trevor Darrell , Yann LeCun

Over the last years, 3D morphable models (3DMMs) have emerged as a state-of-the-art methodology for modeling and generating expressive 3D avatars. However, given their reliance on a strict topology, along with their linear nature, they…

Computer Vision and Pattern Recognition · Computer Science 2025-10-14 Rolandos Alexandros Potamias , Stathis Galanakis , Jiankang Deng , Athanasios Papaioannou , Stefanos Zafeiriou

World models simulate environmental dynamics to enable agents to plan and reason about future states. While existing approaches have primarily focused on visual observations, real-world perception inherently involves multiple sensory…

Multimedia · Computer Science 2026-03-11 Jiahua Wang , Leqi Zheng , Jialong Wu , Yaoxin Mao

Recent interactive video world model methods generate scene evolution conditioned on user instructions. Although they achieve impressive results, two key limitations remain. First, they exhibit motion drift in complex environments with…

Computer Vision and Pattern Recognition · Computer Science 2026-03-19 Guangyuan Li , Bo Li , Jinwei Chen , Xiaobin Hu , Lei Zhao , Peng-Tao Jiang

Learning robot manipulation from abundant human videos offers a scalable alternative to costly robot-specific data collection. However, domain gaps across visual, morphological, and physical aspects hinder direct imitation. To effectively…

Robotics · Computer Science 2025-09-16 Yangcen Liu , Woo Chul Shin , Yunhai Han , Zhenyang Chen , Harish Ravichandar , Danfei Xu

Task-oriented object grasping and rearrangement are critical skills for robots to accomplish different real-world manipulation tasks. However, they remain challenging due to partial observations of the objects and shape variations in…

Robotics · Computer Science 2026-03-06 Yichen Cai , Jianfeng Gao , Christoph Pohl , Tamim Asfour

The field of robotics has made significant strides toward developing generalist robot manipulation policies. However, evaluating these policies in real-world scenarios remains time-consuming and challenging, particularly as the number of…

Robotics · Computer Science 2025-05-27 Yaxuan Li , Yichen Zhu , Junjie Wen , Chaomin Shen , Yi Xu

This paper describes our research on AI agents embodied in visual, virtual or physical forms, enabling them to interact with both users and their environments. These agents, which include virtual avatars, wearable devices, and robots, are…

Deploying generative World-Action Models for manipulation is severely bottlenecked by redundant pixel-level reconstruction, $\mathcal{O}(T)$ memory scaling, and sequential inference latency. We introduce the Causal Latent World Model…

Computer Vision and Pattern Recognition · Computer Science 2026-04-21 Yueci Deng , Guiliang Liu , Kui Jia
‹ Prev 1 3 4 5 6 7 10 Next ›