English
Related papers

Related papers: World Models as Group Actions

200 papers

Humans naturally build mental models of object interactions and dynamics, allowing them to imagine how their surroundings will change if they take a certain action. While generative models today have shown impressive results on…

Computer Vision and Pattern Recognition · Computer Science 2024-08-15 Sruthi Sudhakar , Ruoshi Liu , Basile Van Hoorick , Carl Vondrick , Richard Zemel

Real-time video analysis remains a challenging problem in computer vision, requiring efficient processing of both spatial and temporal information while maintaining computational efficiency. Existing approaches often struggle to balance…

Computer Vision and Pattern Recognition · Computer Science 2025-07-31 Shahla John

Numerous offline and model-based reinforcement learning systems incorporate world models to emulate the inherent environments. A world model is particularly important in scenarios where direct interactions with the real environment is…

Machine Learning · Computer Science 2026-01-19 Rajat Ghosh , Debojyoti Dutta

Assessing action quality is challenging due to the subtle differences between videos and large variations in scores. Most existing approaches tackle this problem by regressing a quality score from a single video, suffering a lot from the…

Computer Vision and Pattern Recognition · Computer Science 2021-08-18 Xumin Yu , Yongming Rao , Wenliang Zhao , Jiwen Lu , Jie Zhou

We propose a method to procedurally generate a familiar yet complex human artifact: the city. We are not trying to reproduce existing cities, but to generate artificial cities that are convincing and plausible by capturing developmental…

Graphics · Computer Science 2025-07-28 Thomas Lechner , Ben Watson , Uri Wilensky , Martin Felsen

Group activity detection (GAD) is the task of identifying members of each group and classifying the activity of the group at the same time in a video. While GAD has been studied recently, there is still much room for improvement in both…

Computer Vision and Pattern Recognition · Computer Science 2024-07-26 Dongkeun Kim , Youngkil Song , Minsu Cho , Suha Kwak

Embodied agents in household environments must plan under partial observation: they need to remember objects, track state changes, and recover when actions fail. Existing benchmarks only partially test this ability. Egocentric video…

Artificial Intelligence · Computer Science 2026-05-14 Qinchuan Cheng , Zhantao Gong , Pengzhan Sun , Angela Yao , Xulei Yang , Shijie Li

Interactive autonomous applications require robustness of the perception engine to artifacts in unconstrained videos. In this paper, we examine the effect of camera motion on the task of action detection. We develop a novel ranking method…

Computer Vision and Pattern Recognition · Computer Science 2022-05-03 Burhan A. Mudassar , Sho Ko , Maojingjing Li , Priyabrata Saha , Saibal Mukhopadhyay

When building a world model, a common assumption is that the environment has a single, unchanging underlying causal rule, like applying Newton's laws to every situation. In reality, what appears as a drifting causal mechanism is often the…

Machine Learning · Computer Science 2025-10-28 Zhiyu Zhao , Haoxuan Li , Haifeng Zhang , Jun Wang , Francesco Faccio , Jürgen Schmidhuber , Mengyue Yang

While a general embodied agent must function as a unified system, current methods are built on isolated models for understanding, world modeling, and control. This fragmentation prevents unifying multimodal generative capabilities and…

Computer Vision and Pattern Recognition · Computer Science 2025-12-29 Hongzhe Bi , Hengkai Tan , Shenghao Xie , Zeyuan Wang , Shuhe Huang , Haitian Liu , Ruowen Zhao , Yao Feng , Chendong Xiang , Yinze Rong , Hongyan Zhao , Hanyu Liu , Zhizhong Su , Lei Ma , Hang Su , Jun Zhu

Inspired by how humans combine direct interaction with action-free experience (e.g., videos), we study world models that learn from heterogeneous data. Standard world models typically rely on action-conditioned trajectories, which limits…

Machine Learning · Computer Science 2025-12-12 Marvin Alles , Xingyuan Zhang , Patrick van der Smagt , Philip Becker-Ehmck

Videos express highly structured spatio-temporal patterns of visual data. A video can be thought of as being governed by two factors: (i) temporally invariant (e.g., person identity), or slowly varying (e.g., activity), attribute-induced…

Computer Vision and Pattern Recognition · Computer Science 2018-03-26 Jiawei He , Andreas Lehrmann , Joseph Marino , Greg Mori , Leonid Sigal

Achieving Artificial General Intelligence (AGI) requires agents that learn and interact adaptively, with interactive world models providing scalable environments for perception, reasoning, and action. Yet current research still lacks…

Computer Vision and Pattern Recognition · Computer Science 2026-05-07 Jianjie Fang , Yingshan Lei , Qin Wan , Ziyou Wang , Yuchao Huang , Yongyan Xu , Baining Zhao , Weichen Zhang , Chen Gao , Xinlei Chen , Yong Li

This paper introduces a probabilistic graphical model for continuous action recognition with two novel components: substructure transition model and discriminative boundary model. The first component encodes the sparse and global temporal…

Computer Vision and Pattern Recognition · Computer Science 2012-03-12 Zhaowen Wang , Jinjun Wang , Jing Xiao , Kai-Hsiang Lin , Thomas Huang

World models have emerged as a critical frontier in AI research, aiming to enhance large models by infusing them with physical dynamics and world knowledge. The core objective is to enable agents to understand, predict, and interact with…

A world model matters to an agent only through the state it constructs. That state must preserve some information, discard other information, and support some future function: prediction, control, planning, memory, grounding, or…

Artificial Intelligence · Computer Science 2026-05-05 Keon Woo Kim

Large-scale video generative models have shown emerging capabilities as zero-shot visual planners, yet video-generated plans often violate temporal consistency and physical constraints, leading to failures when mapped to executable actions.…

Machine Learning · Computer Science 2026-03-17 Christos Ziakas , Amir Bar , Alessandra Russo

Humans have consciousness as the ability to perceive events and objects: a mental model of the world developed from the most impoverished of visual stimuli, enabling humans to make rapid decisions and take actions. Although spatial and…

Artificial Intelligence · Computer Science 2018-11-06 Lisheng Wu , Minne Li , Jun Wang

As AI systems move from generating text to accomplishing goals through sustained interaction, the ability to model environment dynamics becomes a central bottleneck. Agents that manipulate objects, navigate software, coordinate with others,…

State-of-the-art methods for self-supervised sequential action alignment rely on deep networks that find correspondences across videos in time. They either learn frame-to-frame mapping across sequences, which does not leverage temporal…

Computer Vision and Pattern Recognition · Computer Science 2021-11-18 Weizhe Liu , Bugra Tekin , Huseyin Coskun , Vibhav Vineet , Pascal Fua , Marc Pollefeys