中文
相关论文

相关论文: ActionParty: Multi-Subject Action Binding in Gener…

200 篇论文

Training generalist policies for robotic manipulation has shown great promise, as they enable language-conditioned, multi-task behaviors across diverse scenarios. However, evaluating these policies remains difficult because real-world…

机器人学 · 计算机科学 2025-12-05 Wei-Cheng Tseng , Jinwei Gu , Qinsheng Zhang , Hanzi Mao , Ming-Yu Liu , Florian Shkurti , Lin Yen-Chen

We introduce RealPlay, a neural network-based real-world game engine that enables interactive video generation from user control signals. Unlike prior works focused on game-style visuals, RealPlay aims to produce photorealistic, temporally…

计算机视觉与模式识别 · 计算机科学 2025-06-24 Wenqiang Sun , Fangyun Wei , Jinjing Zhao , Xi Chen , Zilong Chen , Hongyang Zhang , Jun Zhang , Yan Lu

A longstanding goal of the field of AI is a method for learning a highly capable, generalist agent from diverse experience. In the subfields of vision and language, this was largely achieved by scaling up transformer-based models and…

Current game world models simulate environments from a subjective, player-centric perspective. However, by treating the Non-Player Character (NPC) merely as background pixels, these models cannot capture interactions between the player and…

计算机视觉与模式识别 · 计算机科学 2026-05-18 Zeqing Wang , Danze Chen , Zhaohu Xing , Zizhao Tong , Yinhan Zhang , Xingyi Yang , Yeying Jin

Autonomous systems have achieved superhuman performance in isolation or simulation, yet they remain brittle in shared, dynamic real-world spaces. This failure stems from the dominant single-agent paradigm for physical applications, where…

机器人学 · 计算机科学 2026-05-22 Ismail Geles , Leonard Bauersfeld , Markus Wulfmeier , Davide Scaramuzza

Anticipating future actions is inherently uncertain. Given an observed video segment containing ongoing actions, multiple subsequent actions can plausibly follow. This uncertainty becomes even larger when predicting far into the future.…

计算机视觉与模式识别 · 计算机科学 2023-11-28 Zeyun Zhong , Chengzhi Wu , Manuel Martin , Michael Voit , Juergen Gall , Jürgen Beyerer

Agents capable of reasoning and planning in the real world require the ability of predicting the consequences of their actions. While world models possess this capability, they most often require action labels, that can be complex to obtain…

人工智能 · 计算机科学 2026-01-21 Quentin Garrido , Tushar Nagarajan , Basile Terver , Nicolas Ballas , Yann LeCun , Michael Rabbat

To go from (passive) process monitoring to active process control, an effective AI system must learn about the behavior of the complex system from very limited training data, forming an ad-hoc digital twin with respect to process inputs and…

As cutting-edge Text-to-Image (T2I) generation models already excel at producing remarkable single images, an even more challenging task, i.e., multi-turn interactive image generation begins to attract the attention of related research…

计算机视觉与模式识别 · 计算机科学 2025-06-02 Junhao Cheng , Xi Lu , Hanhui Li , Khun Loun Zai , Baiqiao Yin , Yuhao Cheng , Yiqiang Yan , Xiaodan Liang

Maintaining consistent characters, props, and environments across multiple shots is a central challenge in narrative video generation. Existing models can produce high-quality short clips but often fail to preserve entity identity and…

计算机视觉与模式识别 · 计算机科学 2026-01-08 Jinsong Zhou , Yihua Du , Xinli Xu , Luozhou Wang , Zijie Zhuang , Yehang Zhang , Shuaibo Li , Xiaojun Hu , Bolan Su , Ying-cong Chen

We present GATSBI, a generative model that can transform a sequence of raw observations into a structured latent representation that fully captures the spatio-temporal context of the agent's actions. In vision-based decision-making…

计算机视觉与模式识别 · 计算机科学 2021-04-12 Cheol-Hui Min , Jinseok Bae , Junho Lee , Young Min Kim

We introduce a framework that enables both multi-view character consistency and 3D camera control in video diffusion models through a novel customization data pipeline. We train the character consistency component with recorded volumetric…

The open-domain video generation models are constrained by the scale of the training video datasets, and some less common actions still cannot be generated. Some researchers explore video editing methods and achieve action generation by…

计算机视觉与模式识别 · 计算机科学 2024-08-26 Jianzhi Liu , Junchen Zhu , Lianli Gao , Heng Tao Shen , Jingkuan Song

Recent video diffusion models have made remarkable strides in visual quality, yet precise, fine-grained control remains a key bottleneck that limits practical customizability for content creation. For AI video creators, three forms of…

计算机视觉与模式识别 · 计算机科学 2026-03-17 Zhenghong Zhou , Xiaohang Zhan , Zhiqin Chen , Soo Ye Kim , Nanxuan Zhao , Haitian Zheng , Qing Liu , He Zhang , Zhe Lin , Yuqian Zhou , Jiebo Luo

In order to enable high-quality decision making and motion planning of intelligent systems such as robotics and autonomous vehicles, accurate probabilistic predictions for surrounding interactive objects is a crucial prerequisite. Although…

机器人学 · 计算机科学 2019-04-05 Jiachen Li , Hengbo Ma , Masayoshi Tomizuka

In multiplayer cooperative video games, players traditionally use individual controllers, inferring others' actions through on-screen visuals and their own movements. This indirect understanding limits truly collaborative gameplay. Research…

Large Action models are essential for enabling autonomous agents to perform complex tasks. However, training such models remains challenging due to the diversity of agent environments and the complexity of noisy agentic data. Existing…

We provide a dataset that enables the creation of learning agents that can build knowledge graph-based world models of interactive narratives. Interactive narratives -- or text-adventure games -- are partially observable environments…

计算与语言 · 计算机科学 2021-06-18 Prithviraj Ammanabrolu , Mark O. Riedl

Video world models have achieved strong visual realism, but this does not ensure that their dynamics are truly governed by actions. In this work, we argue that action faithfulness should be understood through the compositional structure of…

计算机视觉与模式识别 · 计算机科学 2026-05-26 Zijie Wang , Wei Zhang , Weiming Zhang , Fanqi Zhang , Xiao Tan , Yipeng Qin , Guanbin Li

Scalable and reliable evaluation is increasingly critical in the end-to-end era of autonomous driving, where vision--language--action (VLA) policies directly map raw sensor streams to driving actions. Yet, current evaluation pipelines still…

计算机视觉与模式识别 · 计算机科学 2026-04-01 Chaoda Zheng , Sean Li , Jinhao Deng , Zhennan Wang , Shijia Chen , Liqiang Xiao , Ziheng Chi , Hongbin Lin , Kangjie Chen , Boyang Wang , Yu Zhang , Xianming Liu