English
Related papers

Related papers: WildWorld: A Large-Scale Dataset for Dynamic World…

200 papers

In the era of deep learning, data is the critical determining factor in the performance of neural network models. Generating large datasets suffers from various difficulties such as scalability, cost efficiency and photorealism. To avoid…

Computer Vision and Pattern Recognition · Computer Science 2022-10-04 Chahat Deep Singh , Riya Kumari , Cornelia Fermüller , Nitin J. Sanket , Yiannis Aloimonos

World models, especially in autonomous driving, are trending and drawing extensive attention due to their capacity for comprehending driving environments. The established world model holds immense potential for the generation of…

Computer Vision and Pattern Recognition · Computer Science 2023-11-28 Xiaofeng Wang , Zheng Zhu , Guan Huang , Xinze Chen , Jiagang Zhu , Jiwen Lu

Despite the significant impact of visual events on human cognition, understanding events in videos remains a challenging task for AI due to their complex structures, semantic hierarchies, and dynamic evolution. To address this, we propose…

Computer Vision and Pattern Recognition · Computer Science 2025-06-04 Baoyu Liang , Qile Su , Shoutai Zhu , Yuchen Liang , Chao Tong

Robotic manipulation requires anticipating how the environment evolves in response to actions, yet most existing systems lack this predictive capability, often resulting in errors and inefficiency. While Vision-Language Models (VLMs)…

Robotics · Computer Science 2026-02-12 Songen Gu , Yunuo Cai , Tianyu Wang , Simo Wu , Yanwei Fu

Large-scale video generation models have demonstrated emergent physical coherence, positioning them as potential world models. However, a gap remains between contemporary "stateless" video architectures and classic state-centric world model…

Computer Vision and Pattern Recognition · Computer Science 2026-01-27 Luozhou Wang , Zhifei Chen , Yihua Du , Dongyu Yan , Wenhang Ge , Guibao Shen , Xinli Xu , Leyi Wu , Man Chen , Tianshuo Xu , Peiran Ren , Xin Tao , Pengfei Wan , Ying-Cong Chen

This paper presents the World-Action Model (WAM), an action-regularized world model that jointly reasons over future visual observations and the actions that drive state transitions. Unlike conventional world models trained solely via image…

Artificial Intelligence · Computer Science 2026-04-01 Yuci Han , Alper Yilmaz

Adapting pretrained video generation models into controllable world models via latent actions is a promising step towards creating generalist world models. The dominant paradigm adopts a two-stage approach that trains latent action model…

Machine Learning · Computer Science 2026-04-07 Yucen Wang , Fengming Zhang , De-Chuan Zhan , Li Zhao , Kaixin Wang , Jiang Bian

Current world models lack a unified and controlled setting for systematic evaluation, making it difficult to assess whether they truly capture the underlying rules that govern environment dynamics. In this work, we address this open…

Machine Learning · Computer Science 2025-12-01 Xinyi Li , Zaishuo Xia , Weyl Lu , Chenjie Hao , Yubei Chen

Driving world models have gained increasing attention due to their ability to model complex physical dynamics. However, their superb modeling capability is yet to be fully unleashed due to the limited video diversity in current driving…

Computer Vision and Pattern Recognition · Computer Science 2024-10-15 Yuqi Wang , Ke Cheng , Jiawei He , Qitai Wang , Hengchen Dai , Yuntao Chen , Fei Xia , Zhaoxiang Zhang

Foundational world models must be both interactive and preserve spatiotemporal coherence for effective future planning with action choices. However, present models for long video generation have limited inherent world modeling capabilities…

Computer Vision and Pattern Recognition · Computer Science 2026-05-29 Taiye Chen , Xun Hu , Zihan Ding , Chi Jin

A plausible scene evolution depends on the maneuver being considered, while a good maneuver depends on how the scene may evolve. Existing World Action Models (WAMs) largely miss this reciprocity, treating world prediction and action…

Computer Vision and Pattern Recognition · Computer Science 2026-05-13 Hongbo Lu , Liang Yao , Chenghao He , Haoyu Wang , Xiang Gu , Xianfei Li , Wenlong Liao , Tao He , Pai Peng

Inspired by the emergent behaviors in large language models that generalized human intelligence, the research community is pursuing similar emergent capabilities within world models, with a emphasis on modeling the physical world. Within…

Artificial Intelligence · Computer Science 2026-05-21 Kunqi Xu , Jitao Li , Jianglong Ye , Tianshu Tang , Isabella Liu , Sifei Liu , Xueyan Zou

The Driving World Model (DWM), which focuses on predicting scene evolution during the driving process, has emerged as a promising paradigm in the pursuit of autonomous driving (AD). DWMs enable AD systems to better perceive, understand, and…

Computer Vision and Pattern Recognition · Computer Science 2026-02-03 Sifan Tu , Xin Zhou , Dingkang Liang , Xingyu Jiang , Yumeng Zhang , Xiaofan Li , Xiang Bai

This work explores whether a deep generative model can learn complex knowledge solely from visual input, in contrast to the prevalent focus on text-based models like large language models (LLMs). We develop VideoWorld, an auto-regressive…

Computer Vision and Pattern Recognition · Computer Science 2025-03-06 Zhongwei Ren , Yunchao Wei , Xun Guo , Yao Zhao , Bingyi Kang , Jiashi Feng , Xiaojie Jin

This paper introduces Action Image, a new grasp proposal representation that allows learning an end-to-end deep-grasping policy. Our model achieves $84\%$ grasp success on $172$ real world objects while being trained only in simulation on…

Robotics · Computer Science 2020-05-15 Mohi Khansari , Daniel Kappler , Jianlan Luo , Jeff Bingham , Mrinal Kalakrishnan

Interactive world models are advancing rapidly, yet existing benchmarks cover only part of the required competencies, leaving no unified standard for systematic evaluation. To fill this gap, we introduce WBench, a comprehensive multi-turn…

Computer Vision and Pattern Recognition · Computer Science 2026-05-26 Kaining Ying , Hengrui Hu , Siyu Ren , Jiamu Li , Fengjiao Chen , Ziwen Wang , Xuezhi Cao , Xunliang Cai , Henghui Ding

Recent advances in robot foundation models trained on large-scale human teleoperation data have enabled robots to perform increasingly complex real-world tasks. However, scaling these systems remains difficult because collecting…

Multimodal large language models are increasingly deployed as long-horizon agents, where memory must do more than recall: it must track an evolving world, revise what has gone stale, and surface the right evidence at decision time. Existing…

Complex networks serve as abstract models for understanding real-world complex systems and provide frameworks for studying structured dynamical systems. This article addresses limitations in current studies on the exploration of individual…

Social and Information Networks · Computer Science 2025-10-14 Bin Pi , Liang-Jian Deng , Minyu Feng , Matjaž Perc , Jürgen Kurths

Wrist-view observations are crucial for VLA models as they capture fine-grained hand-object interactions that directly enhance manipulation performance. Yet large-scale datasets rarely include such recordings, resulting in a substantial gap…

Computer Vision and Pattern Recognition · Computer Science 2025-10-09 Zezhong Qian , Xiaowei Chi , Yuming Li , Shizun Wang , Zhiyuan Qin , Xiaozhu Ju , Sirui Han , Shanghang Zhang
‹ Prev 1 3 4 5 6 7 10 Next ›