English
Related papers

Related papers: ActionParty: Multi-Subject Action Binding in Gener…

200 papers

Multimodal LLMs are increasingly deployed as perceptual backbones for autonomous agents in 3D environments, from robotics to virtual worlds. These applications require agents to perceive rapid state changes, attribute actions to the correct…

Computation and Language · Computer Science 2026-04-14 Yunzhe Wang , Runhui Xu , Kexin Zheng , Tianyi Zhang , Jayavibhav Niranjan Kogundi , Soham Hans , Volkan Ustun

We argue that 3-D first-person video games are a challenging environment for real-time multi-modal reasoning. We first describe our dataset of human game-play, collected across a large variety of 3-D first-person games, which is both…

Machine Learning · Computer Science 2025-10-21 Yuguang Yue , Irakli Salia , Samuel Hunt , Christopher Green , Wenzhe Shi , Jonathan J Hunt

Action-conditioned video prediction models (often referred to as world models) have shown strong potential for robotics applications, but existing approaches are often slow and struggle to capture physically consistent interactions over…

Action-conditioned video models offer a promising path to building general-purpose robot simulators that can improve directly from data. Yet, despite training on large-scale robot datasets, current state-of-the-art video models still…

The landscape of video generation is shifting, from a focus on generating visually appealing clips to building virtual environments that support interaction and maintain physical plausibility. These developments point toward the emergence…

Artificial Intelligence · Computer Science 2026-02-09 Jingtong Yue , Ziqi Huang , Zhaoxi Chen , Xintao Wang , Pengfei Wan , Ziwei Liu

Recent advances in video diffusion transformers have enabled interactive gaming world models that allow users to explore generated environments over extended horizons. However, existing approaches struggle with precise action control and…

Computer Vision and Pattern Recognition · Computer Science 2026-03-18 Jisu Nam , Yicong Hong , Chun-Hao Paul Huang , Feng Liu , JoungBin Lee , Jiyoung Kim , Siyoon Jin , Yunsung Lee , Jaeyoon Jung , Suhwan Choi , Seungryong Kim , Yang Zhou

Recent successes in autoregressive (AR) generation models, such as the GPT series in natural language processing, have motivated efforts to replicate this success in visual tasks. Some works attempt to extend this approach to autonomous…

Computer Vision and Pattern Recognition · Computer Science 2024-12-31 Xiaotao Hu , Wei Yin , Mingkai Jia , Junyuan Deng , Xiaoyang Guo , Qian Zhang , Xiaoxiao Long , Ping Tan

We present Playable Environments - a new representation for interactive video generation and manipulation in space and time. With a single image at inference time, our novel framework allows the user to move objects in 3D while generating a…

Computer Vision and Pattern Recognition · Computer Science 2022-03-17 Willi Menapace , Stéphane Lathuilière , Aliaksandr Siarohin , Christian Theobalt , Sergey Tulyakov , Vladislav Golyanik , Elisa Ricci

Generative models trained on internet data have revolutionized how text, image, and video content can be created. Perhaps the next milestone for generative models is to simulate realistic experience in response to actions taken by humans,…

Artificial Intelligence · Computer Science 2024-09-27 Sherry Yang , Yilun Du , Kamyar Ghasemipour , Jonathan Tompson , Leslie Kaelbling , Dale Schuurmans , Pieter Abbeel

We present a novel study on enhancing the capability of preserving the content in world models, focusing on a property we term World Stability. Recent diffusion-based generative models have advanced the synthesis of immersive and realistic…

Machine Learning · Computer Science 2025-03-12 Soonwoo Kwon , Jin-Young Kim , Hyojun Go , Kyungjune Baek

World modeling is a crucial task for enabling intelligent agents to effectively interact with humans and operate in dynamic environments. In this work, we propose MineWorld, a real-time interactive world model on Minecraft, an open-ended…

Computer Vision and Pattern Recognition · Computer Science 2025-04-14 Junliang Guo , Yang Ye , Tianyu He , Haoyu Wu , Yushu Jiang , Tim Pearce , Jiang Bian

We introduce Genie, the first generative interactive environment trained in an unsupervised manner from unlabelled Internet videos. The model can be prompted to generate an endless variety of action-controllable virtual worlds described…

Spatio-temporal action detection is an important and challenging problem in video understanding. The existing action detection benchmarks are limited in aspects of small numbers of instances in a trimmed video or low-level atomic actions.…

Computer Vision and Pattern Recognition · Computer Science 2021-08-19 Yixuan Li , Lei Chen , Runyu He , Zhenzhi Wang , Gangshan Wu , Limin Wang

Extracting the rules of real-world multi-agent behaviors is a current challenge in various scientific and engineering fields. Biological agents independently have limited observation and mechanical constraints; however, most of the…

Machine Learning · Computer Science 2023-12-04 Keisuke Fujii , Naoya Takeishi , Yoshinobu Kawahara , Kazuya Takeda

Recent advances in world models have greatly enhanced interactive environment simulation. Existing methods mainly fall into two categories: (1) static world generation models, which construct 3D environments without active agents, and (2)…

Computer Vision and Pattern Recognition · Computer Science 2025-12-22 Yitong Wang , Fangyun Wei , Hongyang Zhang , Bo Dai , Yan Lu

World action models (WAMs) have emerged as a promising direction for robot policy learning, as they can leverage powerful video backbones to model the future states. However, existing approaches often rely on separate action modules, or use…

Computer Vision and Pattern Recognition · Computer Science 2026-04-16 Haoyu Zhen , Zixian Gao , Qiao Sun , Yilin Zhao , Yuncong Yang , Yilun Du , Pengsheng Guo , Tsun-Hsuan Wang , Yi-Ling Qiao , Chuang Gan

Our world is full of varied actions and moves across specialized domains that we, as humans, strive to identify and understand. Within any single domain, actions can often appear quite similar, making it challenging for deep models to…

Computer Vision and Pattern Recognition · Computer Science 2024-11-12 Mohammadreza Salehi , Jae Sung Park , Tanush Yadav , Aditya Kusupati , Ranjay Krishna , Yejin Choi , Hannaneh Hajishirzi , Ali Farhadi

Video generation models, as one form of world models, have emerged as one of the most exciting frontiers in AI, promising agents the ability to imagine the future by modeling the temporal evolution of complex scenes. In autonomous driving,…

Computer Vision and Pattern Recognition · Computer Science 2026-03-10 Yang Zhou , Hao Shao , Letian Wang , Zhuofan Zong , Hongsheng Li , Steven L. Waslander

Generating customized content in videos has received increasing attention recently. However, existing works primarily focus on customized text-to-video generation for single subject, suffering from subject-missing and attribute-binding…

Computer Vision and Pattern Recognition · Computer Science 2024-05-22 Hong Chen , Xin Wang , Yipeng Zhang , Yuwei Zhou , Zeyang Zhang , Siao Tang , Wenwu Zhu

Video-based world models have recently garnered increasing attention for their ability to synthesize diverse and dynamic visual environments. In this paper, we focus on shared world modeling, where a model generates multiple videos from a…

Computer Vision and Pattern Recognition · Computer Science 2025-12-03 Fan Wu , Jiacheng Wei , Ruibo Li , Yi Xu , Junyou Li , Deheng Ye , Guosheng Lin