English
Related papers

Related papers: iMoWM: Taming Interactive Multi-Modal World Model …

200 papers

Recent Multimodal Large Language Models (MLLMs) have shown high potential for spatial reasoning within 3D scenes. However, they typically rely on computationally expensive 3D representations like point clouds or reconstructed Bird's-Eye…

Computer Vision and Pattern Recognition · Computer Science 2026-05-08 Shuyao Shi , Kang G. Shin

Recent research has been increasingly focusing on developing 3D world models that simulate complex real-world scenarios. World models have found broad applications across various domains, including embodied AI, autonomous driving,…

Artificial Intelligence · Computer Science 2025-09-09 Yinglin Duan , Zhengxia Zou , Tongwei Gu , Wei Jia , Zhan Zhao , Luyi Xu , Xinzhu Liu , Yenan Lin , Hao Jiang , Kang Chen , Shuang Qiu

World models enable robots to conduct counterfactual reasoning in physical environments by predicting future world states. While conventional approaches often prioritize pixel-level reconstruction of future scenes, such detailed rendering…

Robotics · Computer Science 2025-12-22 Zhiwei Zhang , Hui Zhang , Kaihong Huang , Chenghao Shi , Huimin Lu

Humans understand the world through the integration of multiple sensory modalities, enabling them to perceive, reason about, and imagine dynamic physical processes. Inspired by this capability, multimodal foundation models (MFMs) have…

Artificial Intelligence · Computer Science 2025-10-07 Xuehai He

The ability to simulate the effects of future actions on the world is a crucial ability of intelligent embodied agents, enabling agents to anticipate the effects of their actions and make plans accordingly. While a large body of existing…

Computer Vision and Pattern Recognition · Computer Science 2025-05-12 Siyuan Zhou , Yilun Du , Yuncong Yang , Lei Han , Peihao Chen , Dit-Yan Yeung , Chuang Gan

Recent video diffusion foundation models have achieved remarkable progress in high-quality video generation, yet turning them into real-time interactive video world models remains challenging. Interactive world models require controllable,…

Computer Vision and Pattern Recognition · Computer Science 2026-05-29 Min Zhao , Hongzhou Zhu , Bokai Yan , Zihan Zhou , Yimin Chen , Wenqiang Sun , Kaiwen Zheng , Guande He , Xiao Yang , Chongxuan Li , Fan Bao , Jun Zhu

Traditional 3D morphable face models (3DMMs) provide fine-grained control over expression but cannot easily capture geometric and appearance details. Neural volumetric representations approach photorealism but are hard to animate and do not…

Computer Vision and Pattern Recognition · Computer Science 2022-11-07 Yufeng Zheng , Victoria Fernández Abrevaya , Marcel C. Bühler , Xu Chen , Michael J. Black , Otmar Hilliges

Large Language Models (LLMs) have gained popularity in task planning for long-horizon manipulation tasks. To enhance the validity of LLM-generated plans, visual demonstrations and online videos have been widely employed to guide the…

Robotics · Computer Science 2025-03-12 Kejia Chen , Zheng Shen , Yue Zhang , Lingyun Chen , Fan Wu , Zhenshan Bing , Sami Haddadin , Alois Knoll

Legged locomotion over various terrains is challenging and requires precise perception of the robot and its surroundings from both proprioception and vision. However, learning directly from high-dimensional visual input is often…

Robotics · Computer Science 2024-09-26 Hang Lai , Jiahang Cao , Jiafeng Xu , Hongtao Wu , Yunfeng Lin , Tao Kong , Yong Yu , Weinan Zhang

Enabling embodied agents to imagine future states is essential for robust and generalizable visual navigation. Yet, state-of-the-art systems typically rely on modular designs that decouple navigation planning from visual world modeling,…

Artificial Intelligence · Computer Science 2026-03-24 Yifei Dong , Fengyi Wu , Guangyu Chen , Lingdong Kong , Xu Zhu , Qiyu Hu , Yuxuan Zhou , Jingdong Sun , Jun-Yan He , Qi Dai , Alexander G. Hauptmann , Zhi-Qi Cheng

World models have recently re-emerged as a central paradigm for embodied intelligence, robotics, autonomous driving, and model-based reinforcement learning. However, current world model research is often dominated by three partially…

Artificial Intelligence · Computer Science 2026-05-27 Sen Cui , Jingheng Ma

Learning predictive models from high-dimensional sensory observations is fundamental for cyber-physical systems, yet the latent representations learned by standard world models lack physical interpretability. This limits their reliability,…

Machine Learning · Computer Science 2026-04-07 Zhenjiang Mao , Mrinall Eashaan Umasudhan , Ivan Ruchkin

Embodied world models have emerged as a promising paradigm in robotics, most of which leverage large-scale Internet videos or pretrained video generation models to enrich visual and motion priors. However, they still face key challenges: a…

Robotics · Computer Science 2026-02-04 Yixiang Chen , Peiyan Li , Jiabing Yang , Keji He , Xiangnan Wu , Yuan Xu , Kai Wang , Jing Liu , Nianfeng Liu , Yan Huang , Liang Wang

Identifying predictive world models for robots in novel environments from sparse online observations is essential for robot task planning and execution in novel environments. However, existing methods that leverage differentiable…

Robotics · Computer Science 2025-05-13 Yifan Zhu , Tianyi Xiang , Aaron Dollar , Zherong Pan

World models aim to endow AI systems with the ability to represent, generate, and interact with dynamic environments in a coherent and temporally consistent manner. While recent video generation models have demonstrated impressive visual…

Accurate modeling and simulation of mobile networks are essential for enabling intelligent and cost-effective network optimization. In this paper, we propose MobiWorld, a generative world model designed to support high-fidelity and flexible…

Networking and Internet Architecture · Computer Science 2025-07-15 Haoye Chai , Yuan Yuan , Yong Li

World Action Models (WAMs) enhance Vision-Language-Action policies by jointly predicting scene evolution and robot actions, but existing methods usually represent the predicted world as holistic images, video tokens, or global latents.…

While non-prehensile manipulation (e.g., controlled pushing/poking) constitutes a foundational robotic skill, its learning remains challenging due to the high sensitivity to complex physical interactions involving friction and restitution.…

Machine Learning · Computer Science 2025-05-06 Wenxuan Li , Hang Zhao , Zhiyuan Yu , Yu Du , Qin Zou , Ruizhen Hu , Kai Xu

Masked image modeling (MIM) has emerged as a promising approach for pre-training Vision Transformers (ViTs). MIMs predict masked tokens token-wise to recover target signals that are tokenized from images or generated by pre-trained models…

Computer Vision and Pattern Recognition · Computer Science 2025-03-24 Taekyung Kim , Byeongho Heo , Dongyoon Han

Robotic world models are a promising paradigm for forecasting future environment states, yet their inference speed and the physical plausibility of generated trajectories remain critical bottlenecks, limiting their real-world applications.…

Robotics · Computer Science 2025-09-26 Sibo Li , Qianyue Hao , Yu Shang , Yong Li