English
Related papers

Related papers: Grounding Video Models to Actions through Goal Con…

200 papers

A practical navigation agent must be capable of handling a wide range of interaction demands, such as following instructions, searching objects, answering questions, tracking people, and more. Existing models for embodied navigation fall…

How can artificial agents learn to solve many diverse tasks in complex visual environments in the absence of any supervision? We decompose this question into two problems: discovering new goals and learning to reliably achieve them. We…

Machine Learning · Computer Science 2021-10-19 Russell Mendonca , Oleh Rybkin , Kostas Daniilidis , Danijar Hafner , Deepak Pathak

Complex, real-world domains may not be fully modeled for an agent, especially if the agent has never operated in the domain before. The agent's ability to effectively plan and act in such a domain is influenced by its knowledge of when it…

Artificial Intelligence · Computer Science 2022-03-08 Dustin Dannenhauer , Matthew Molineaux , Michael W. Floyd , Noah Reifsnyder , David W. Aha

This paper presents an unsupervised approach towards automatically extracting video-based guidance on object usage, from egocentric video and wearable gaze tracking, collected from multiple users while performing tasks. The approach i)…

Computer Vision and Pattern Recognition · Computer Science 2016-03-22 Dima Damen , Teesid Leelasawassuk , Walterio Mayol-Cuevas

Despite tremendous progress in dexterous manipulation, current visuomotor policies remain fundamentally limited by two challenges: they struggle to generalize under perceptual or behavioral distribution shifts, and their performance is…

Robotics · Computer Science 2025-08-04 Junbang Liang , Pavel Tokmakov , Ruoshi Liu , Sruthi Sudhakar , Paarth Shah , Rares Ambrus , Carl Vondrick

Simple as it seems, moving an object to another location within an image is, in fact, a challenging image-editing task that requires re-harmonizing the lighting, adjusting the pose based on perspective, accurately filling occluded regions,…

Graphics · Computer Science 2025-03-12 Xin Yu , Tianyu Wang , Soo Ye Kim , Paul Guerrero , Xi Chen , Qing Liu , Zhe Lin , Xiaojuan Qi

Camera-controlled video generation has achieved remarkable progress in recent years. However, existing video-to-video re-rendering methods primarily rely on Supervised Fine-Tuning using synthetic datasets. At present, there is an extreme…

Computer Vision and Pattern Recognition · Computer Science 2026-05-25 Zizun Li , Haoyu Guo , Runzhe Teng , Chunhua Shen , Tong He

Action Detection is a complex task that aims to detect and classify human actions in video clips. Typically, it has been addressed by processing fine-grained features extracted from a video classification backbone. Recently, thanks to the…

Computer Vision and Pattern Recognition · Computer Science 2021-03-02 Matteo Tomei , Lorenzo Baraldi , Simone Calderara , Simone Bronzin , Rita Cucchiara

Egocentric world models present a promising direction for enabling agents to predict and plan, but their performance is constrained by the limited availability of egocentric training data and its inherent partial observability of humans'…

Computer Vision and Pattern Recognition · Computer Science 2026-05-27 Danny Tran , Roberto Martín-Martín , Kristen Grauman

Video world models have achieved strong visual realism, but this does not ensure that their dynamics are truly governed by actions. In this work, we argue that action faithfulness should be understood through the compositional structure of…

Computer Vision and Pattern Recognition · Computer Science 2026-05-26 Zijie Wang , Wei Zhang , Weiming Zhang , Fanqi Zhang , Xiao Tan , Yipeng Qin , Guanbin Li

Recent vision-language-action (VLA) models rely on 2D inputs, lacking integration with the broader realm of the 3D physical world. Furthermore, they perform action prediction by learning a direct mapping from perception to action,…

Computer Vision and Pattern Recognition · Computer Science 2024-03-15 Haoyu Zhen , Xiaowen Qiu , Peihao Chen , Jincheng Yang , Xin Yan , Yilun Du , Yining Hong , Chuang Gan

World models simulate future states of the world in response to different actions. They facilitate interactive content creation and provides a foundation for grounded, long-horizon reasoning. Current foundation models do not fully meet the…

Computer Vision and Pattern Recognition · Computer Science 2024-06-17 Jiannan Xiang , Guangyi Liu , Yi Gu , Qiyue Gao , Yuting Ning , Yuheng Zha , Zeyu Feng , Tianhua Tao , Shibo Hao , Yemin Shi , Zhengzhong Liu , Eric P. Xing , Zhiting Hu

A major bottleneck for developing general reinforcement learning agents is determining rewards that will yield desirable behaviors under various circumstances. We introduce a general mechanism for automatically specifying meaningful…

Machine Learning · Computer Science 2017-11-22 Ashley D. Edwards , Charles L. Isbell

We address the problem of generating a 3D-consistent, navigable environment that is spatially grounded: a simulation of a real location. Existing video generative models can produce a plausible sequence that is consistent with a text (T2V)…

Computer Vision and Pattern Recognition · Computer Science 2026-04-22 Gene Chou , Charles Herrmann , Kyle Genova , Boyang Deng , Songyou Peng , Bharath Hariharan , Jason Y. Zhang , Noah Snavely , Philipp Henzler

We introduce environment predictive coding, a self-supervised approach to learn environment-level representations for embodied agents. In contrast to prior work on self-supervised learning for images, we aim to jointly encode a series of…

Computer Vision and Pattern Recognition · Computer Science 2021-02-05 Santhosh K. Ramakrishnan , Tushar Nagarajan , Ziad Al-Halah , Kristen Grauman

The rapid development of large language and multimodal models has sparked significant interest in using proprietary models, such as GPT-4o, to develop autonomous agents capable of handling real-world scenarios like web navigation. Although…

Computation and Language · Computer Science 2024-10-28 Hongliang He , Wenlin Yao , Kaixin Ma , Wenhao Yu , Hongming Zhang , Tianqing Fang , Zhenzhong Lan , Dong Yu

Future robots are envisioned as versatile systems capable of performing a variety of household tasks. The big question remains, how can we bridge the embodiment gap while minimizing physical robot learning, which fundamentally does not…

Robotics · Computer Science 2025-03-31 Hanzhi Chen , Boyang Sun , Anran Zhang , Marc Pollefeys , Stefan Leutenegger

Video generation has been used to generate visual plans for controlling robotic systems. Given an image observation and a language instruction, previous work has generated video plans which are then converted to robot controls to be…

Artificial Intelligence · Computer Science 2025-02-11 Achint Soni , Sreyas Venkataraman , Abhranil Chandra , Sebastian Fischmeister , Percy Liang , Bo Dai , Sherry Yang

We introduce a general framework for visual forecasting, which directly imitates visual sequences without additional supervision. As a result, our model can be applied at several semantic levels and does not require any domain knowledge or…

Computer Vision and Pattern Recognition · Computer Science 2017-08-22 Kuo-Hao Zeng , William B. Shen , De-An Huang , Min Sun , Juan Carlos Niebles

Autonomy is a hallmark of animal intelligence, enabling adaptive and intelligent behavior in complex environments without relying on external reward or task structure. Existing reinforcement learning approaches to exploration in reward-free…

Neurons and Cognition · Quantitative Biology 2025-10-27 Reece Keller , Alyn Kirsch , Felix Pei , Xaq Pitkow , Leo Kozachkov , Aran Nayebi