English
Related papers

Related papers: World-in-World: World Models in a Closed-Loop Worl…

200 papers

Video world models have achieved strong visual realism, but this does not ensure that their dynamics are truly governed by actions. In this work, we argue that action faithfulness should be understood through the compositional structure of…

Computer Vision and Pattern Recognition · Computer Science 2026-05-26 Zijie Wang , Wei Zhang , Weiming Zhang , Fanqi Zhang , Xiao Tan , Yipeng Qin , Guanbin Li

While video-generation-based embodied world models have gained increasing attention, their reliance on large-scale embodied interaction data remains a key bottleneck. The scarcity, difficulty of collection, and high dimensionality of…

Computer Vision and Pattern Recognition · Computer Science 2026-03-30 Hao Li , Qiao Sun

Training generalist policies for robotic manipulation has shown great promise, as they enable language-conditioned, multi-task behaviors across diverse scenarios. However, evaluating these policies remains difficult because real-world…

Robotics · Computer Science 2025-12-05 Wei-Cheng Tseng , Jinwei Gu , Qinsheng Zhang , Hanzi Mao , Ming-Yu Liu , Florian Shkurti , Lin Yen-Chen

Closed-loop evaluation of autonomous-driving policies requires interactive simulation beyond log replay. However, existing generative world models often degrade in closed loop due to (i) history-free initialization that mismatches policy…

Robotics · Computer Science 2026-03-19 Chaokang Jiang , Desen Zhou , Jiuming Liu , Kevin Li Sun

Video-generative world models are increasingly used as neural simulators for embodied planning and policy learning, yet their ability to predict physical risk and severe consequences is rarely evaluated.We find that these models often…

Robotics · Computer Science 2026-04-21 Zhenglin Lai , Sirui Huang , Yuteng Li , Changxin Huang , Jianqiang Li , Bingzhe Wu

As Vision-Language Models (VLMs) are increasingly deployed as autonomous cognitive cores for embodied assistants, evaluating their privacy awareness in physical environments becomes critical. Unlike digital chatbots, these agents operate in…

Cryptography and Security · Computer Science 2026-05-11 Junran Wang , Xinjie Shen , Zehao Jin , Pan Li

We present a novel study on enhancing the capability of preserving the content in world models, focusing on a property we term World Stability. Recent diffusion-based generative models have advanced the synthesis of immersive and realistic…

Machine Learning · Computer Science 2025-03-12 Soonwoo Kwon , Jin-Young Kim , Hyojun Go , Kyungjune Baek

Training robotic policies directly in the real world is expensive and unscalable. Although generative simulation enables large-scale data synthesis, current approaches often fail to generate logically coherent long-horizon tasks and…

Robotics · Computer Science 2026-02-13 Xiang Liu , Sen Cui , Guocai Yao , Zhong Cao , Jingheng Ma , Min Zhang , Changshui Zhang

The imitation learning of self-driving vehicle policies through behavioral cloning is often carried out in an open-loop fashion, ignoring the effect of actions to future states. Training such policies purely with Empirical Risk Minimization…

Machine Learning · Computer Science 2022-10-12 Eesha Kumar , Yiming Zhang , Stefano Pini , Simon Stent , Ana Ferreira , Sergey Zagoruyko , Christian S. Perone

Recent advances in agent development have focused on scaling model size and raw interaction data, mirroring successes in large language models. However, for complex, long-horizon multi-agent tasks such as robotic soccer, this end-to-end…

Artificial Intelligence · Computer Science 2025-11-05 Brennen Hill

Scalable Embodied AI faces fundamental constraints due to prohibitive costs and safety risks of real-world interaction. While Embodied World Models (EWMs) offer promise through imagined rollouts, existing approaches suffer from geometric…

Computer Vision and Pattern Recognition · Computer Science 2026-04-14 Ruicheng Zhang , Guangyu Chen , Zunnan Xu , Zihao Liu , Zhizhou Zhong , Mingyang Zhang , Jun Zhou , Xiu Li

Designing adaptive mechanisms to align individual and collective interests remains a central challenge in artificial social intelligence. Existing methods often struggle with modeling heterogeneous agents possessing persistent latent traits…

Computers and Society · Computer Science 2025-10-23 Xiaoyuan Zhang , Yizhe Huang , Chengdong Ma , Zhixun Chen , Long Ma , Yali Du , Song-Chun Zhu , Yaodong Yang , Xue Feng

World-Action Models (WAM) initialized from pre-trained video generation backbones have demonstrated remarkable potential for robot policy learning. However, existing approaches face two critical bottlenecks that hinder performance and…

World generation is a fundamental capability for applications like video games, simulation, and robotics. However, existing approaches face three main obstacles: controllability, scalability, and efficiency. End-to-end scene generation…

Computer Vision and Pattern Recognition · Computer Science 2026-01-28 Han-Hung Lee , Cheng-Yu Yang , Yu-Lun Liu , Angel X. Chang

The ability to perceive and reason about individual objects and their interactions is a goal to be achieved for building intelligent artificial systems. State-of-the-art approaches use a feedforward encoder to extract object embeddings and…

Computer Vision and Pattern Recognition · Computer Science 2024-02-07 Jonathan Collu , Riccardo Majellaro , Aske Plaat , Thomas M. Moerland

Video world models have attracted significant attention for their ability to produce high-fidelity future visual observations conditioned on past observations and navigation actions. Temporally- and spatially-consistent, long-term world…

Computer Vision and Pattern Recognition · Computer Science 2025-12-03 Yuta Oshima , Yusuke Iwasawa , Masahiro Suzuki , Yutaka Matsuo , Hiroki Furuta

Learned world models hold significant potential for robotic manipulation, as they can serve as simulator for real-world interactions. While extensive progress has been made in 2D video-based world models, these approaches often lack…

Robotics · Computer Science 2025-10-13 Chuanrui Zhang , Zhengxian Wu , Guanxing Lu , Yansong Tang , Ziwei Wang

Pretrained video generation models provide strong priors for robot control, but existing unified world action models still struggle to decode reliable actions without substantial robot-specific training. We attribute this limitation to a…

Robotics · Computer Science 2026-04-14 Liaoyuan Fan , Zetian Xu , Chen Cao , Wenyao Zhang , Mingqi Yuan , Jiayu Chen

Video--based world models have emerged along two dominant paradigms: video generation and 3D reconstruction. However, existing evaluation benchmarks either focus narrowly on visual fidelity and text--video alignment for generative models,…

Computer Vision and Pattern Recognition · Computer Science 2026-03-24 Meiqi Wu , Zhixin Cai , Fufangchen Zhao , Xiaokun Feng , Rujing Dang , Bingze Song , Ruitian Tian , Jiashu Zhu , Jiachen Lei , Hao Dou , Jing Tang , Lei Sun , Jiahong Wu , Xiangxiang Chu , Zeming Liu , Kaiqi Huang
‹ Prev 1 8 9 10 Next ›