English
Related papers

Related papers: ActionParty: Multi-Subject Action Binding in Gener…

200 papers

Training generalist policies for robotic manipulation has shown great promise, as they enable language-conditioned, multi-task behaviors across diverse scenarios. However, evaluating these policies remains difficult because real-world…

Robotics · Computer Science 2025-12-05 Wei-Cheng Tseng , Jinwei Gu , Qinsheng Zhang , Hanzi Mao , Ming-Yu Liu , Florian Shkurti , Lin Yen-Chen

We introduce RealPlay, a neural network-based real-world game engine that enables interactive video generation from user control signals. Unlike prior works focused on game-style visuals, RealPlay aims to produce photorealistic, temporally…

Computer Vision and Pattern Recognition · Computer Science 2025-06-24 Wenqiang Sun , Fangyun Wei , Jinjing Zhao , Xi Chen , Zilong Chen , Hongyang Zhang , Jun Zhang , Yan Lu

A longstanding goal of the field of AI is a method for learning a highly capable, generalist agent from diverse experience. In the subfields of vision and language, this was largely achieved by scaling up transformer-based models and…

Current game world models simulate environments from a subjective, player-centric perspective. However, by treating the Non-Player Character (NPC) merely as background pixels, these models cannot capture interactions between the player and…

Computer Vision and Pattern Recognition · Computer Science 2026-05-18 Zeqing Wang , Danze Chen , Zhaohu Xing , Zizhao Tong , Yinhan Zhang , Xingyi Yang , Yeying Jin

Autonomous systems have achieved superhuman performance in isolation or simulation, yet they remain brittle in shared, dynamic real-world spaces. This failure stems from the dominant single-agent paradigm for physical applications, where…

Robotics · Computer Science 2026-05-22 Ismail Geles , Leonard Bauersfeld , Markus Wulfmeier , Davide Scaramuzza

Anticipating future actions is inherently uncertain. Given an observed video segment containing ongoing actions, multiple subsequent actions can plausibly follow. This uncertainty becomes even larger when predicting far into the future.…

Computer Vision and Pattern Recognition · Computer Science 2023-11-28 Zeyun Zhong , Chengzhi Wu , Manuel Martin , Michael Voit , Juergen Gall , Jürgen Beyerer

Agents capable of reasoning and planning in the real world require the ability of predicting the consequences of their actions. While world models possess this capability, they most often require action labels, that can be complex to obtain…

Artificial Intelligence · Computer Science 2026-01-21 Quentin Garrido , Tushar Nagarajan , Basile Terver , Nicolas Ballas , Yann LeCun , Michael Rabbat

To go from (passive) process monitoring to active process control, an effective AI system must learn about the behavior of the complex system from very limited training data, forming an ad-hoc digital twin with respect to process inputs and…

As cutting-edge Text-to-Image (T2I) generation models already excel at producing remarkable single images, an even more challenging task, i.e., multi-turn interactive image generation begins to attract the attention of related research…

Computer Vision and Pattern Recognition · Computer Science 2025-06-02 Junhao Cheng , Xi Lu , Hanhui Li , Khun Loun Zai , Baiqiao Yin , Yuhao Cheng , Yiqiang Yan , Xiaodan Liang

Maintaining consistent characters, props, and environments across multiple shots is a central challenge in narrative video generation. Existing models can produce high-quality short clips but often fail to preserve entity identity and…

Computer Vision and Pattern Recognition · Computer Science 2026-01-08 Jinsong Zhou , Yihua Du , Xinli Xu , Luozhou Wang , Zijie Zhuang , Yehang Zhang , Shuaibo Li , Xiaojun Hu , Bolan Su , Ying-cong Chen

We present GATSBI, a generative model that can transform a sequence of raw observations into a structured latent representation that fully captures the spatio-temporal context of the agent's actions. In vision-based decision-making…

Computer Vision and Pattern Recognition · Computer Science 2021-04-12 Cheol-Hui Min , Jinseok Bae , Junho Lee , Young Min Kim

We introduce a framework that enables both multi-view character consistency and 3D camera control in video diffusion models through a novel customization data pipeline. We train the character consistency component with recorded volumetric…

Computer Vision and Pattern Recognition · Computer Science 2025-10-17 Yuancheng Xu , Wenqi Xian , Li Ma , Julien Philip , Ahmet Levent Taşel , Yiwei Zhao , Ryan Burgert , Mingming He , Oliver Hermann , Oliver Pilarski , Rahul Garg , Paul Debevec , Ning Yu

The open-domain video generation models are constrained by the scale of the training video datasets, and some less common actions still cannot be generated. Some researchers explore video editing methods and achieve action generation by…

Computer Vision and Pattern Recognition · Computer Science 2024-08-26 Jianzhi Liu , Junchen Zhu , Lianli Gao , Heng Tao Shen , Jingkuan Song

Recent video diffusion models have made remarkable strides in visual quality, yet precise, fine-grained control remains a key bottleneck that limits practical customizability for content creation. For AI video creators, three forms of…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Zhenghong Zhou , Xiaohang Zhan , Zhiqin Chen , Soo Ye Kim , Nanxuan Zhao , Haitian Zheng , Qing Liu , He Zhang , Zhe Lin , Yuqian Zhou , Jiebo Luo

In order to enable high-quality decision making and motion planning of intelligent systems such as robotics and autonomous vehicles, accurate probabilistic predictions for surrounding interactive objects is a crucial prerequisite. Although…

Robotics · Computer Science 2019-04-05 Jiachen Li , Hengbo Ma , Masayoshi Tomizuka

In multiplayer cooperative video games, players traditionally use individual controllers, inferring others' actions through on-screen visuals and their own movements. This indirect understanding limits truly collaborative gameplay. Research…

Human-Computer Interaction · Computer Science 2024-11-19 Kenta Hashiura , Kazuya Iida , Takeru Hashimoto , Youichi Kamiyama , Keita Watanabe , Kouta Minamizawa , Takuji Narumi

Large Action models are essential for enabling autonomous agents to perform complex tasks. However, training such models remains challenging due to the diversity of agent environments and the complexity of noisy agentic data. Existing…

We provide a dataset that enables the creation of learning agents that can build knowledge graph-based world models of interactive narratives. Interactive narratives -- or text-adventure games -- are partially observable environments…

Computation and Language · Computer Science 2021-06-18 Prithviraj Ammanabrolu , Mark O. Riedl

Video world models have achieved strong visual realism, but this does not ensure that their dynamics are truly governed by actions. In this work, we argue that action faithfulness should be understood through the compositional structure of…

Computer Vision and Pattern Recognition · Computer Science 2026-05-26 Zijie Wang , Wei Zhang , Weiming Zhang , Fanqi Zhang , Xiao Tan , Yipeng Qin , Guanbin Li

Scalable and reliable evaluation is increasingly critical in the end-to-end era of autonomous driving, where vision--language--action (VLA) policies directly map raw sensor streams to driving actions. Yet, current evaluation pipelines still…

Computer Vision and Pattern Recognition · Computer Science 2026-04-01 Chaoda Zheng , Sean Li , Jinhao Deng , Zhennan Wang , Shijia Chen , Liqiang Xiao , Ziheng Chi , Hongbin Lin , Kangjie Chen , Boyang Wang , Yu Zhang , Xianming Liu