English
Related papers

Related papers: EVA: An Embodied World Model for Future Video Anti…

200 papers

Trained on internet-scale video data, generative world models are increasingly recognized as powerful world simulators that can generate consistent and plausible dynamics over structure, motion, and physics. This raises a natural question:…

Computer Vision and Pattern Recognition · Computer Science 2025-10-02 Kevin Zhang , Kuangzhi Ge , Xiaowei Chi , Renrui Zhang , Shaojun Shi , Zhen Dong , Sirui Han , Shanghang Zhang

Humans can perceive and reason about spatial relationships from sequential visual observations, such as egocentric video streams. However, how pretrained models acquire such abilities, especially high-level reasoning, remains unclear. This…

Artificial Intelligence · Computer Science 2025-04-18 Baining Zhao , Ziyou Wang , Jianjie Fang , Chen Gao , Fanhang Man , Jinqiang Cui , Xin Wang , Xinlei Chen , Yong Li , Wenwu Zhu

Video generation models, as one form of world models, have emerged as one of the most exciting frontiers in AI, promising agents the ability to imagine the future by modeling the temporal evolution of complex scenes. In autonomous driving,…

Computer Vision and Pattern Recognition · Computer Science 2026-03-10 Yang Zhou , Hao Shao , Letian Wang , Zhuofan Zong , Hongsheng Li , Steven L. Waslander

Foundational world models must be both interactive and preserve spatiotemporal coherence for effective future planning with action choices. However, present models for long video generation have limited inherent world modeling capabilities…

Computer Vision and Pattern Recognition · Computer Science 2026-05-29 Taiye Chen , Xun Hu , Zihan Ding , Chi Jin

In real-world deployment, vision-language models often encounter disturbances such as weather, occlusion, and camera motion. Under such conditions, their understanding and reasoning degrade substantially, revealing a gap between clean,…

Computer Vision and Pattern Recognition · Computer Science 2026-04-15 Yangfan He , Changgyu Boo , Jaehong Yoon

As world models gain momentum in Embodied AI, an increasing number of works explore using video foundation models as predictive world models for downstream embodied tasks like 3D prediction or interactive generation. However, before…

Recent advances in creative AI have enabled the synthesis of high-fidelity images and videos conditioned on language instructions. Building on these developments, text-to-video diffusion models have evolved into embodied world models (EWMs)…

Robotics · Computer Science 2025-05-20 Hu Yue , Siyuan Huang , Yue Liao , Shengcong Chen , Pengfei Zhou , Liliang Chen , Maoqing Yao , Guanghui Ren

Embodied visual planning aims to enable manipulation tasks by imagining how a scene evolves toward a desired goal and using the imagined trajectories to guide actions. Video diffusion models, through their image-to-video generation…

Computer Vision and Pattern Recognition · Computer Science 2025-12-30 Yuming Gu , Yizhi Wang , Yining Hong , Yipeng Gao , Hao Jiang , Angtian Wang , Bo Liu , Nathaniel S. Dennler , Zhengfei Kuang , Hao Li , Gordon Wetzstein , Chongyang Ma

World models, which are predictive representations of how environments evolve under actions, have become a central component of robot learning. They support policy learning, planning, simulation, evaluation, data generation, and have…

Achieving human-like reasoning in deep learning models for complex tasks in unknown environments remains a critical challenge in embodied intelligence. While advanced vision-language models (VLMs) excel in static scene understanding, their…

Computer Vision and Pattern Recognition · Computer Science 2025-09-03 Jinzhou Tang , Jusheng zhang , Sidi Liu , Waikit Xiu , Qinhan Lv , Xiying Li

Video generation models have emerged as high-fidelity models of the physical world, capable of synthesizing high-quality videos capturing fine-grained interactions between agents and their environments conditioned on multi-modal user…

While language models have become impactful in many real-world applications, video generation remains largely confined to entertainment. Motivated by video's inherent capacity to demonstrate physical-world information that is difficult to…

Computer Vision and Pattern Recognition · Computer Science 2025-11-25 Junhao Cheng , Liang Hou , Xin Tao , Jing Liao

Leveraging future observation modeling to facilitate action generation presents a promising avenue for enhancing the capabilities of Vision-Language-Action (VLA) models. However, existing approaches struggle to strike a balance between…

The ability to simulate the effects of future actions on the world is a crucial ability of intelligent embodied agents, enabling agents to anticipate the effects of their actions and make plans accordingly. While a large body of existing…

Computer Vision and Pattern Recognition · Computer Science 2025-05-12 Siyuan Zhou , Yilun Du , Yuncong Yang , Lei Han , Peihao Chen , Dit-Yan Yeung , Chuang Gan

Achieving general-purpose robotics requires empowering robots to adapt and evolve based on their environment and feedback. Traditional methods face limitations such as extensive training requirements, difficulties in cross-task…

Robotics · Computer Science 2026-04-23 Jianzong Wang , Botao Zhao , Yayun He , Junqing Peng , Xulong Zhang

World models are widely explored in embodied intelligence, yet they typically predict distinct evolutions of the world and the ego within a single stream, where the world captures persistent instruction-agnostic scene regularities and the…

Computer Vision and Pattern Recognition · Computer Science 2026-05-20 Zuyao Lin , Jianhui Zhang , Peidong Jia , Xiaoguang Zhao , Shanghang Zhang , Xingyu Chen

While Multimodal Large Language Models demonstrate impressive semantic capabilities, they often suffer from spatial blindness, struggling with fine-grained geometric reasoning and physical dynamics. Existing solutions typically rely on…

Computer Vision and Pattern Recognition · Computer Science 2026-03-20 Xianjin Wu , Dingkang Liang , Tianrui Feng , Kui Xia , Yumeng Zhang , Xiaofan Li , Xiao Tan , Xiang Bai

To effectively engage in human society, the ability to adapt, filter information, and make informed decisions in ever-changing situations is critical. As robots and intelligent agents become more integrated into human life, there is a…

Artificial Intelligence · Computer Science 2025-11-13 Mingyang Mao , Mariela M. Perez-Cabarcas , Utteja Kallakuri , Nicholas R. Waytowich , Xiaomin Lin , Tinoosh Mohsenin

Recent large vision-language models have achieved strong performance on short- and medium-length video understanding, yet they remain inadequate for ultra-long or even infinite video reasoning, where models must preserve coherent memory…

Artificial Intelligence · Computer Science 2026-05-08 Peizheng Yan , Yu Zhao , Liang Xie , Juntong Qi , Mingming Wang , Erwei Yin

Video reasoning models are a core component of egocentric and embodied agents. However, standard benchmarks for assessing models provide only evaluation of the output (e.g. the answer to a question), without evaluation of intermediate…

Computer Vision and Pattern Recognition · Computer Science 2026-05-18 Arsha Nagrani , Jasper Uijilings , Shyamal Buch , Tobias Weyand , Sudheendra Vijayanarasimhan , Bo Hu , Ramin Mehran , David A Ross , Cordelia Schmid