English
Related papers

Related papers: iMoWM: Taming Interactive Multi-Modal World Model …

200 papers

The ability to perform complex tasks from detailed instructions is a key to many remarkable achievements of our species. As humans, we are not only capable of performing a wide variety of tasks but also very complex ones that may entail…

Artificial Intelligence · Computer Science 2024-07-23 Xiaoxuan Lei , Lucas Gomez , Hao Yuan Bai , Pouya Bashivan

Multimodal Language Language Models (MLLMs) demonstrate the emerging abilities of "world models" -- interpreting and reasoning about complex real-world dynamics. To assess these abilities, we posit videos are the ideal medium, as they…

Computer Vision and Pattern Recognition · Computer Science 2024-07-31 Xuehai He , Weixi Feng , Kaizhi Zheng , Yujie Lu , Wanrong Zhu , Jiachen Li , Yue Fan , Jianfeng Wang , Linjie Li , Zhengyuan Yang , Kevin Lin , William Yang Wang , Lijuan Wang , Xin Eric Wang

World modeling has become a cornerstone in AI research, enabling agents to understand, represent, and predict the dynamic environments they inhabit. While prior work largely emphasizes generative methods for 2D image and video data, they…

Large-scale video generative models can synthesize diverse and realistic visual content for dynamic world creation, but they often lack element-wise controllability, hindering their use in editing scenes and training embodied AI agents. We…

Computer Vision and Pattern Recognition · Computer Science 2025-06-10 Sicheng Mo , Ziyang Leng , Leon Liu , Weizhen Wang , Honglin He , Bolei Zhou

Scalable Embodied AI faces fundamental constraints due to prohibitive costs and safety risks of real-world interaction. While Embodied World Models (EWMs) offer promise through imagined rollouts, existing approaches suffer from geometric…

Computer Vision and Pattern Recognition · Computer Science 2026-04-14 Ruicheng Zhang , Guangyu Chen , Zunnan Xu , Zihao Liu , Zhizhou Zhong , Mingyang Zhang , Jun Zhou , Xiu Li

Learning robust and scalable visual representations from massive multi-view video data remains a challenge in computer vision and autonomous driving. Existing pre-training methods either rely on expensive supervised learning with 3D…

Computer Vision and Pattern Recognition · Computer Science 2024-03-14 Jialv Zou , Bencheng Liao , Qian Zhang , Wenyu Liu , Xinggang Wang

Mobile traffic prediction is a fundamental yet challenging problem for wireless network planning and optimization. Existing models focus on learning static long-term temporal patterns in mobile traffic series, which limits their ability to…

Networking and Internet Architecture · Computer Science 2026-04-10 Xiaoqian Qi , Haoye Chai , Yue Wang , Yong Li

World models based on video generation demonstrate remarkable potential for simulating interactive environments but face persistent difficulties in two key areas: maintaining long-term content consistency when scenes are revisited and…

Computer Vision and Pattern Recognition · Computer Science 2026-02-27 Tianxing Xu , Zixuan Wang , Guangyuan Wang , Li Hu , Zhongyi Zhang , Peng Zhang , Bang Zhang , Song-Hai Zhang

Learning-based 3D object reconstruction enables single- or few-shot estimation of 3D object models. For robotics, this holds the potential to allow model-based methods to rapidly adapt to novel objects and scenes. Existing 3D reconstruction…

Understanding and replicating the real world is a critical challenge in Artificial General Intelligence (AGI) research. To achieve this, many existing approaches, such as world models, aim to capture the fundamental principles governing the…

Computer Vision and Pattern Recognition · Computer Science 2026-02-17 Yuqi Hu , Longguang Wang , Xian Liu , Ling-Hao Chen , Yuwei Guo , Yukai Shi , Ce Liu , Anyi Rao , Zeyu Wang , Hui Xiong

Recent advances in creative AI have enabled the synthesis of high-fidelity images and videos conditioned on language instructions. Building on these developments, text-to-video diffusion models have evolved into embodied world models (EWMs)…

Robotics · Computer Science 2025-05-20 Hu Yue , Siyuan Huang , Yue Liao , Shengcong Chen , Pengfei Zhou , Liliang Chen , Maoqing Yao , Guanghui Ren

World models (WMs) represent the frontier of sample-efficient reinforcement learning, but their complexity leaves many promising improvements unrealized due to the significant expertise and effort required to identify and integrate them.…

Machine Learning · Computer Science 2026-05-12 Lior Cohen , Kaixin Wang , Bingyi Kang , Uri Gadot , Shie Mannor

We introduce Robowheel, a data engine that converts human hand object interaction (HOI) videos into training-ready supervision for cross morphology robotic learning. From monocular RGB or RGB-D inputs, we perform high precision HOI…

Despite encouraging progress in 3D scene understanding, it remains challenging to develop an effective Large Multi-modal Model (LMM) that is capable of understanding and reasoning in complex 3D environments. Most previous methods typically…

Computer Vision and Pattern Recognition · Computer Science 2025-06-17 Hanxun Yu , Wentong Li , Song Wang , Junbo Chen , Jianke Zhu

World models are progressively being employed across diverse fields, extending from basic environment simulation to complex scenario construction. However, existing models are mainly trained on domain-specific states and actions, and…

Artificial Intelligence · Computer Science 2024-10-01 Zhiqi Ge , Hongzhe Huang , Mingze Zhou , Juncheng Li , Guoming Wang , Siliang Tang , Yueting Zhuang

Recent advances in control robot methods, from end-to-end vision-language-action frameworks to modular systems with predefined primitives, have advanced robots' ability to follow natural language instructions. Nonetheless, many approaches…

Building robots that can automate labor-intensive tasks has long been the core motivation behind the advancements in computer vision and the robotics community. Recent interest in leveraging 3D algorithms, particularly neural fields, has…

Computer Vision and Pattern Recognition · Computer Science 2023-12-13 Litian Liang , Liuyu Bian , Caiwei Xiao , Jialin Zhang , Linghao Chen , Isabella Liu , Fanbo Xiang , Zhiao Huang , Hao Su

This work focuses on generating realistic, physically-based human behaviors from multi-modal inputs, which may only partially specify the desired motion. For example, the input may come from a VR controller providing arm motion and body…

Robotics · Computer Science 2025-02-11 Aayam Shrestha , Pan Liu , German Ros , Kai Yuan , Alan Fern

Recent advancements in robotics have enabled robots to navigate complex scenes or manipulate diverse objects independently. However, robots are still impotent in many household tasks requiring coordinated behaviors such as opening doors.…

Robotics · Computer Science 2024-12-09 Ruihan Yang , Yejin Kim , Rose Hendrix , Aniruddha Kembhavi , Xiaolong Wang , Kiana Ehsani

Imitation learning from large-scale, diverse human demonstrations has been shown to be effective for training robots, but collecting such data is costly and time-consuming. This challenge intensifies for multi-step bimanual mobile…