English
Related papers

Related papers: Action Images: End-to-End Policy Learning via Mult…

200 papers

Video world models have achieved remarkable success in simulating environmental dynamics in response to actions by users or agents. They are modeled as action-conditioned video generation models that take historical frames and current…

Computer Vision and Pattern Recognition · Computer Science 2026-04-22 Haoyu Wu , Jiwen Yu , Yingtian Zou , Xihui Liu

World models predict future transitions from observations and actions. Existing works predominantly focus on image generation only. Visual feature-based world models, on the other hand, predict future visual features instead of raw video…

Computer Vision and Pattern Recognition · Computer Science 2026-05-11 Xinyu Zhang , Zhengtong Xu , Yutian Tao , Yeping Wang , Yu She , Abdeslam Boularias

Embodied AI requires agents that perceive, act, and anticipate how actions reshape future world states. World models serve as internal simulators that capture environment dynamics, enabling forward and counterfactual rollouts to support…

Computer Vision and Pattern Recognition · Computer Science 2025-12-02 Xinqing Li , Xin He , Le Zhang , Min Wu , Xiaoli Li , Yun Liu

Generating robot motion that fulfills multiple tasks simultaneously is challenging due to the geometric constraints imposed by the robot. In this paper, we propose to solve multi-task problems through learning structured policies from human…

Robotics · Computer Science 2021-03-12 M. Asif Rana , Anqi Li , Dieter Fox , Sonia Chernova , Byron Boots , Nathan Ratliff

World models derived from large-scale video generative pre-training have emerged as a promising paradigm for generalist robot policy learning. However, standard approaches often focus on high-fidelity RGB video prediction, this can result…

World models, which are predictive representations of how environments evolve under actions, have become a central component of robot learning. They support policy learning, planning, simulation, evaluation, data generation, and have…

Video generation models are rapidly improving in their ability to synthesize human actions in novel contexts, holding the potential to serve as high-level planners for contextual robot control. To realize this potential, a key research…

Robotics · Computer Science 2025-12-12 James Ni , Zekai Wang , Wei Lin , Amir Bar , Yann LeCun , Trevor Darrell , Jitendra Malik , Roei Herzig

Conventional visuomotor imitation learning usually predicts future robot actions directly in the time domain. Such formulations often have limited physical scene awareness and weak long-horizon memory. In contrast, world-model-based…

Robotics · Computer Science 2026-05-12 Changchuan Yang , Yuhang Dong , Guanzhong Tian , Haizhou Ge , Hongrui Zhu

In recent years, video action recognition, as a fundamental task in the field of video understanding, has been deeply explored by numerous researchers.Most traditional video action recognition methods typically involve converting videos…

Computer Vision and Pattern Recognition · Computer Science 2024-08-20 Junlin Chen , Chengcheng Xu , Yangfan Xu , Jian Yang , Jun Li , Zhiping Shi

Pretrained video diffusion models provide powerful spatiotemporal generative priors, making them a natural foundation for robotic world models. While recent world-action models jointly optimize future videos and actions, they predominantly…

Computer Vision and Pattern Recognition · Computer Science 2026-05-08 Zhaoyang Yang , Yurun Jin , Lizhe Qi , Cong Huang , Kai Chen

Action recognition is a fundamental ability for social species. Yet, its underlying computations are not well understood. Classical psychophysical studies using simplified stimuli have shown that humans can perceive body motion even under…

Computer Vision and Pattern Recognition · Computer Science 2026-04-21 Prerana Kumar , Martin A. Giese

The goal of this paper is to recognize actions in video without the need for examples. Different from traditional zero-shot approaches we do not demand the design and specification of attribute classifiers and class-to-attribute mappings to…

Computer Vision and Pattern Recognition · Computer Science 2015-10-26 Mihir Jain , Jan C. van Gemert , Thomas Mensink , Cees G. M. Snoek

Dynamical systems theory and reinforcement learning view world evolution as latent-state dynamics driven by actions, with visual observations providing partial information about the state. Recent video world models attempt to learn this…

Computer Vision and Pattern Recognition · Computer Science 2026-03-25 Zhen Li , Zian Meng , Shuwei Shi , Wenshuo Peng , Yuwei Wu , Bo Zheng , Chuanhao Li , Kaipeng Zhang

Robot action planning in the real world is challenging as it requires not only understanding the current state of the environment but also predicting how it will evolve in response to actions. Vision-language-action (VLA), which repurpose…

We introduce the Do-Undo task and benchmark to address a critical gap in vision-language models: understanding and generating plausible scene transformations driven by real-world actions. Unlike prior work that relies on prompt-based image…

Computer Vision and Pattern Recognition · Computer Science 2026-05-15 Shweta Mahajan , Shreya Kadambi , Hoang Le , Rajeev Yasarla , Apratim Bhattacharyya , Munawar Hayat , Fatih Porikli

Collision-free motion generation in unknown environments is a core building block for robot manipulation. Generating such motions is challenging due to multiple objectives; not only should the solutions be optimal, the motion generator…

Robotics · Computer Science 2022-10-25 Adam Fishman , Adithyavairan Murali , Clemens Eppner , Bryan Peele , Byron Boots , Dieter Fox

Recent advances in video diffusion transformers have enabled interactive gaming world models that allow users to explore generated environments over extended horizons. However, existing approaches struggle with precise action control and…

Computer Vision and Pattern Recognition · Computer Science 2026-03-18 Jisu Nam , Yicong Hong , Chun-Hao Paul Huang , Feng Liu , JoungBin Lee , Jiyoung Kim , Siyoon Jin , Yunsung Lee , Jaeyoon Jung , Suhwan Choi , Seungryong Kim , Yang Zhou

Control policy learning for modular robot locomotion has previously been limited to proprioceptive feedback and flat terrain. This paper develops policies for modular systems with vision traversing more challenging environments. These…

Robotics · Computer Science 2023-05-02 Julian Whitman , Howie Choset

Reasoning is central to purposeful action, yet most robotic foundation models map perception and instructions directly to control, which limits adaptability, generalization, and semantic grounding. We introduce Action Reasoning Models…

General-purpose world models promise scalable policy evaluation, optimization, and planning, yet achieving the required level of robustness remains challenging. Unlike policy learning, which primarily focuses on optimal actions, a world…

Machine Learning · Computer Science 2026-04-03 Yuejiang Liu , Fan Feng , Lingjing Kong , Weifeng Lu , Jinzhou Tang , Kun Zhang , Kevin Murphy , Chelsea Finn , Yilun Du