English
Related papers

Related papers: AIM: Intent-Aware Unified world action Modeling wi…

200 papers

Human actions manipulating articulated objects, such as opening and closing a drawer, can be categorized into multiple modalities we define as interaction modes. Traditional robot learning approaches lack discrete representations of these…

Robotics · Computer Science 2024-10-29 Liquan Wang , Ankit Goyal , Haoping Xu , Animesh Garg

Recent advancements in vision-language-action (VLA) models have shown promise in robotic manipulation, yet they continue to struggle with long-horizon, multi-step tasks. Existing methods lack internal reasoning mechanisms that can identify…

Visual prediction has emerged as a promising paradigm for embodied control, where future observations are generated and then translated into actions. However, dense video generation is computationally expensive and often unnecessary for…

Computer Vision and Pattern Recognition · Computer Science 2026-05-20 Yiren Song , Yihan Wang , Xiyao Deng , Zhuoran Yan , Mike Zheng Shou

Visual Robot Manipulation (VRM) aims to enable a robot to follow natural language instructions based on robot states and visual observations, and therefore requires costly multi-modal data. To compensate for the deficiency of robot data,…

Computer Vision and Pattern Recognition · Computer Science 2025-08-12 Dejie Yang , Zijing Zhao , Yang Liu

Mental simulation is a critical cognitive function for goal-directed behavior because it is essential for assessing actions and their consequences. When a self-generated or externally specified goal is given, a sequence of actions that is…

Robotics · Computer Science 2019-03-13 Minju Jung , Takazumi Matsumoto , Jun Tani

Reward design remains a critical bottleneck in visual reinforcement learning (RL) for robotic manipulation. In simulated environments, rewards are conventionally designed based on the distance to a target position. However, such precise…

Machine Learning · Computer Science 2025-09-29 Nan Tang , Jing-Cheng Pang , Guanlin Li , Chao Qian , Yang Yu

Learning-based methods are promising to plan robot motion without performing extensive search, which is needed by many non-learning approaches. Recently, Value Iteration Networks (VINs) received much interest since---in contrast to standard…

Robotics · Computer Science 2019-07-02 Daniel Schleich , Tobias Klamt , Sven Behnke

Enabling robots to explore and act in unfamiliar environments under ambiguous human instructions by interactively identifying task-relevant objects (e.g., identifying cups or beverages for "I'm thirsty") remains challenging for existing…

Robotics · Computer Science 2026-02-06 Hengxuan Xu , Fengbo Lan , Zhixin Zhao , Shengjie Wang , Mengqiao Liu , Jieqian Sun , Yu Cheng , Tao Zhang

While end-to-end autonomous driving has achieved remarkable progress in geometric control, current systems remain constrained by a command-following paradigm that relies on simple navigational instructions. Transitioning to genuinely…

Computer Vision and Pattern Recognition · Computer Science 2026-01-08 Huan Zheng , Yucheng Zhou , Tianyi Yan , Jiayi Su , Hongjun Chen , Dubing Chen , Xingtai Gui , Wencheng Han , Runzhou Tao , Zhongying Qiu , Jianfei Yang , Jianbing Shen

Typical end-to-end formulations for learning robotic navigation involve predicting a small set of steering command actions (e.g., step forward, turn left, turn right, etc.) from images of the current state (e.g., a bird's-eye view of a SLAM…

Robotics · Computer Science 2020-10-13 Jimmy Wu , Xingyuan Sun , Andy Zeng , Shuran Song , Johnny Lee , Szymon Rusinkiewicz , Thomas Funkhouser

End-to-end autonomous driving systems increasingly rely on vision-centric world models to understand and predict their environment. However, a common ineffectiveness in these models is the full reconstruction of future scenes, which expends…

Computer Vision and Pattern Recognition · Computer Science 2026-02-10 Jianbiao Mei , Yu Yang , Xuemeng Yang , Licheng Wen , Jiajun Lv , Botian Shi , Yong Liu

A robot's instantaneous sensory observations do not always reveal task-relevant state information. Under such partial observability, optimal behavior typically involves explicitly acting to gain the missing information. Today's standard…

It has been observed that deep neural networks (DNNs) often use both genuine as well as spurious features. In this work, we propose "Amending Inherent Interpretability via Self-Supervised Masking" (AIM), a simple yet interestingly effective…

Computer Vision and Pattern Recognition · Computer Science 2025-08-18 Eyad Alshami , Shashank Agnihotri , Bernt Schiele , Margret Keuper

Embodied navigation in open, dynamic environments demands accurate foresight of how the world will evolve and how actions will unfold over time. We propose AstraNav-World, an end-to-end world model that jointly reasons about future visual…

Computer Vision and Pattern Recognition · Computer Science 2026-04-09 Jintao Chen , Junjun Hu , Haochen Bai , Minghua Luo , Xinda Xue , Botao Ren , Chengyu Bai , Shichao Xie , Ziyi Chen , Fei Liu , Zedong Chu , Xiaolong Wu , Mu Xu , Shanghang Zhang

With the goal of increasing the speed and efficiency in robotic manipulation, a control approach is presented that aims to utilize intentional simultaneous impacts to its advantage. This approach exploits the concept of the time-invariant…

Robotics · Computer Science 2024-11-18 Jari van Steen , Nathan van de Wouw , Alessandro Saccon

A critical gap exists between the general-purpose visual understanding of state-of-the-art physical AI models and the specialized perceptual demands of structured real-world deployment environments. We present PRISM, a 270K-sample…

Computer Vision and Pattern Recognition · Computer Science 2026-04-01 Amirreza Rouhi , Parikshit Sakurikar , Satya Sai Reddy , Narsimha Menga , Anirudh Govil , Sri Harsha Chittajallu , Rajat Aggarwal , Anoop Namboodiri , Sashi Reddi

Action understanding, encompassing action detection and anticipation, plays a crucial role in numerous practical applications. However, untrimmed videos are often characterized by substantial redundant information and noise. Moreover, in…

Computer Vision and Pattern Recognition · Computer Science 2026-02-24 Xinyu Yang , Zheheng Jiang , Feixiang Zhou , Yihang Zhu , Na Lv , Nan Xing , Nishan Canagarajah , Huiyu Zhou

Current efforts to learn scalable policies in robotic manipulation primarily fall into two categories: one focuses on "action," which involves behavior cloning from extensive collections of robotic data, while the other emphasizes "vision,"…

Robotics · Computer Science 2024-12-20 Yang Tian , Sizhe Yang , Jia Zeng , Ping Wang , Dahua Lin , Hao Dong , Jiangmiao Pang

Motion planning involves determining a sequence of robot configurations to reach a desired pose, subject to movement and safety constraints. Traditional motion planning finds collision-free paths, but this is overly restrictive in clutter,…

Robotics · Computer Science 2026-03-10 Yiyang Ling , Karan Owalekar , Oluwatobiloba Adesanya , Erdem Bıyık , Daniel Seita

Generating talking avatars is a fundamental task in video generation. Although existing methods can generate full-body talking avatars with simple human motion, extending this task to grounded human-object interaction (GHOI) remains an open…

Computer Vision and Pattern Recognition · Computer Science 2026-02-04 Youliang Zhang , Zhengguang Zhou , Zhentao Yu , Ziyao Huang , Teng Hu , Sen Liang , Guozhen Zhang , Ziqiao Peng , Shunkai Li , Yi Chen , Zixiang Zhou , Yuan Zhou , Qinglin Lu , Xiu Li