English
Related papers

Related papers: GazeVLA: Learning Human Intention for Robotic Mani…

200 papers

Imitation learning for acquiring generalizable policies often requires a large volume of demonstration data, making the process significantly costly. One promising strategy to address this challenge is to leverage the cognitive and…

Robotics · Computer Science 2025-06-09 Yutaro Ishida , Takamitsu Matsubara , Takayuki Kanai , Kazuhiro Shintani , Hiroshi Bito

Human gaze is known to be a strong indicator of underlying human intentions and goals during manipulation tasks. This work studies gaze patterns of human teachers demonstrating tasks to robots and proposes ways in which such patterns can be…

Robotics · Computer Science 2021-11-30 Akanksha Saran , Elaine Schaertl Short , Andrea Thomaz , Scott Niekum

The study of human-robot interaction is fundamental to the design and use of robotics in real-world applications. Robots will need to predict and adapt to the actions of human collaborators in order to achieve good performance and improve…

In collaborative human-robot manipulation, a robot must predict human intents and adapt its actions accordingly to smoothly execute tasks. However, the human's intent in turn depends on actions the robot takes, creating a chicken-or-egg…

Robotics · Computer Science 2024-06-04 Kushal Kedia , Atiksh Bhardwaj , Prithwish Dan , Sanjiban Choudhury

Human intention-based systems enable robots to perceive and interpret user actions to interact with humans and adapt to their behavior proactively. Therefore, intention prediction is pivotal in creating a natural interaction with social…

Robotics · Computer Science 2025-04-09 Hassan Ali , Philipp Allgeuer , Stefan Wermter

Robot imitation data are often multimodal: similar visual-language observations may be followed by different action chunks because human demonstrators act with different short-horizon intents, task phases, or recent context. Existing…

Predicting human motion is critical for assistive robots and AR/VR applications, where the interaction with humans needs to be safe and comfortable. Meanwhile, an accurate prediction depends on understanding both the scene context and human…

Computer Vision and Pattern Recognition · Computer Science 2022-07-20 Yang Zheng , Yanchao Yang , Kaichun Mo , Jiaman Li , Tao Yu , Yebin Liu , C. Karen Liu , Leonidas J. Guibas

People are proficient at communicating their intentions in order to avoid conflicts when navigating in narrow, crowded environments. In many situations mobile robots lack both the ability to interpret human intentions and the ability to…

Robotics · Computer Science 2019-11-07 Justin Hart , Reuth Mirsky , Stone Tejeda , Bonny Mahajan , Jamin Goo , Kathryn Baldauf , Sydney Owen , Peter Stone

Intention recognition, or the ability to anticipate the actions of another agent, plays a vital role in the design and development of automated assistants that can support humans in their daily tasks. In particular, industrial settings pose…

Artificial Intelligence · Computer Science 2024-11-27 Juan Carlos Saborio , Joachim Hertzberg

Predicting future human pose is a fundamental application for machine intelligence, which drives robots to plan their behavior and paths ahead of time to seamlessly accomplish human-robot collaboration in real-world 3D scenarios. Despite…

Computer Vision and Pattern Recognition · Computer Science 2024-05-07 Zhenyu Lou , Qiongjie Cui , Haofan Wang , Xu Tang , Hong Zhou

In gaze based Human-Robot Interaction (HRI), it is important to determine the human intention for further interaction. The gaze intention is often modelled as fixation. However, when looking at an object, it is not natural and it is…

Robotics · Computer Science 2019-09-18 Lei Shi , Cosmin Copot , Steve Vanlanduit

We propose a novel system for robot-to-human object handover that emulates human coworker interactions. Unlike most existing studies that focus primarily on grasping strategies and motion planning, our system focus on 1. inferring human…

Robotics · Computer Science 2025-03-06 Hanxin Zhang , Abdulqader Dhafer , Zhou Daniel Hao , Hongbiao Dong

The overarching goal of this work is to efficiently enable end-users to correctly anticipate a robot's behavior in novel situations. Since a robot's behavior is often a direct result of its underlying objective function, our insight is that…

Robotics · Computer Science 2018-10-19 Sandy H. Huang , David Held , Pieter Abbeel , Anca D. Dragan

Eye movements have long been studied as a window into the attentional mechanisms of the human brain and made accessible as novelty style human-machine interfaces. However, not everything that we gaze upon, is something we want to interact…

Computer Vision and Pattern Recognition · Computer Science 2022-01-25 Paul Festor , Ali Shafti , Alex Harston , Michey Li , Pavel Orlov , A. Aldo Faisal

In computer vision, video-based approaches have been widely explored for the early classification and the prediction of actions or activities. However, it remains unclear whether this modality (as compared to 3D kinematics) can still be…

Computer Vision and Pattern Recognition · Computer Science 2017-08-04 Andrea Zunino , Jacopo Cavazza , Atesh Koul , Andrea Cavallo , Cristina Becchio , Vittorio Murino

Manipulation tasks in daily life, such as pouring water, unfold intentionally under specialized manipulation contexts. Being able to process contextual knowledge in these Activities of Daily Living (ADLs) over time can help us understand…

Computer Vision and Pattern Recognition · Computer Science 2020-03-04 Chen Jiang , Masood Dehghan , Martin Jagersand

Forecasting future 3D hand pose sequences from egocentric video is essential for understanding human intention and enabling embodied applications such as AR/VR assistance and human-robot interaction. However, this task remains a highly…

Computer Vision and Pattern Recognition · Computer Science 2026-05-11 Jaeyoung Choi , Hyeondong Kim , Yujin Kim , Daehee Park

Vision-language-action (VLA) models perform well on training-seen robotic tasks but struggle to generalize to unseen scenes and objects. A key limitation lies in their implicit visual representations, which entangle object appearance,…

Computer Vision and Pattern Recognition · Computer Science 2026-05-08 Hanyu Zhou , Chuanhao Ma , Gim Hee Lee

Recent advances in Vision-Language-Action (VLA) models have enabled robotic agents to integrate multimodal understanding with action execution. However, our empirical analysis reveals that current VLAs struggle to allocate visual attention…

This work aims to tackle the intent recognition problem in Human-Robot Collaborative assembly scenarios. Precisely, we consider an interactive assembly of a wooden stool where the robot fetches the pieces in the correct order and the human…