English
Related papers

Related papers: Intent at a Glance: Gaze-Guided Robotic Manipulati…

200 papers

The human gaze is an important cue to signal intention, attention, distraction, and the regions of interest in the immediate surroundings. Gaze tracking can transform how robots perceive, understand, and react to people, enabling new modes…

Robotics · Computer Science 2024-06-11 Tim Schreiter , Andrey Rudenko , Martin Magnusson , Achim J. Lilienthal

Eye-in-hand cameras have shown promise in enabling greater sample efficiency and generalization in vision-based robotic manipulation. However, for robotic imitation, it is still expensive to have a human teleoperator collect large amounts…

Robotics · Computer Science 2023-07-13 Moo Jin Kim , Jiajun Wu , Chelsea Finn

Large Language Models (LLMs) have substantially improved the conversational capabilities of social robots. Nevertheless, for an intuitive and fluent human-robot interaction, robots should be able to ground the conversation by relating…

Human-Computer Interaction · Computer Science 2026-04-09 Elisabeth Menendez , Michael Gienger , Santiago Martínez , Carlos Balaguer , Anna Belardinelli

Mobile manipulation is a fundamental capability for general-purpose robotic agents, requiring both coordinated control of the mobile base and manipulator and robust perception under dynamically changing viewpoints. However, existing…

Robotics · Computer Science 2026-04-30 Jiahao Liu , Cui Wenbo , Zhongpu Xia , Haoran Li , Dongbin Zhao

Robotic affordances, providing information about what actions can be taken in a given situation, can aid robotic manipulation. However, learning about affordances requires expensive large annotated datasets of interactions or…

Robotics · Computer Science 2024-06-07 Pietro Mazzaglia , Taco Cohen , Daniel Dijkman

Robotic affordances, providing information about what actions can be taken in a given situation, can aid robotic manipulation. However, learning about affordances requires expensive large annotated datasets of interactions or…

Robotics · Computer Science 2024-06-14 Pietro Mazzaglia , Taco Cohen , Daniel Dijkman

World models allow autonomous agents to plan and explore by predicting the visual outcomes of different actions. However, for robot manipulation, it is challenging to accurately model the fine-grained robot-object interaction within the…

Robotics · Computer Science 2025-07-30 Fangqi Zhu , Hongtao Wu , Song Guo , Yuxiao Liu , Chilam Cheang , Tao Kong

In dynamic environments such as warehouses, hospitals, and homes, robots must seamlessly transition between gross motion and precise manipulations to complete complex tasks. However, current Vision-Language-Action (VLA) frameworks, largely…

Performing robotic grasping from a cluttered bin based on human instructions is a challenging task, as it requires understanding both the nuances of free-form language and the spatial relationships between objects. Vision-Language Models…

We introduce CARMA, a system for situational grounding in human-robot group interactions. Effective collaboration in such group settings requires situational awareness based on a consistent representation of present persons and objects…

Traditional approaches to training agents have generally involved a single, deterministic environment of minimal complexity to solve various tasks such as robot locomotion or computer vision. However, agents trained in static environments…

Robotics · Computer Science 2025-10-01 Kevin Godin-Dubois , Karine Miras , Anna V. Kononova

Vision-Language-Action (VLA) models leverage pretrained vision-language models (VLMs) to couple perception with robotic control, offering a promising path toward general-purpose embodied intelligence. However, current SOTA VLAs are…

Robotics · Computer Science 2025-10-10 Yandu Chen , Kefan Gu , Yuqing Wen , Yucheng Zhao , Tiancai Wang , Liqiang Nie

Object-goal visual navigation requires robots to reason over semantic structure and act effectively under partial observability. Recent approaches based on object-level topological maps enable long-horizon navigation without dense geometric…

Robotics · Computer Science 2026-03-27 Yanmei Jiao , Anpeng Lu , Wenhan Hu , Rong Xiong , Yue Wang , Huajin Tang , Wen-an Zhang

Routine inspections for critical infrastructures such as bridges are required in most jurisdictions worldwide. Such routine inspections are largely visual in nature, which are qualitative, subjective, and not repeatable. Although robotic…

Robotics · Computer Science 2024-08-12 Sunwoong Choi , Zaid Abbas Al-Sabbag , Sriram Narasimhan , Chul Min Yeum

This study investigates the potential of eye-tracking technology and the Segment Anything Model (SAM) to design a collaborative human-computer interaction system that automates medical image segmentation. We present the \textbf{GazeSAM}…

Computer Vision and Pattern Recognition · Computer Science 2023-04-28 Bin Wang , Armstrong Aboah , Zheyuan Zhang , Ulas Bagci

Gaze following aims to interpret human-scene interactions by predicting the person's focal point of gaze. Prevailing approaches often adopt a two-stage framework, whereby multi-modality information is extracted in the initial stage for gaze…

Computer Vision and Pattern Recognition · Computer Science 2024-11-15 Yuehao Song , Xinggang Wang , Jingfeng Yao , Wenyu Liu , Jinglin Zhang , Xiangmin Xu

Collaborative manipulation is inherently multimodal, with haptic communication playing a central role. When performed by humans, it involves back-and-forth force exchanges between the participants through which they resolve possible…

Robotics · Computer Science 2023-08-21 Zhanibek Rysbek , Ki Hwan Oh , Milos Zefran

Within this work, we explore intention inference for user actions in the context of a handheld robot setup. Handheld robots share the shape and properties of handheld tools while being able to process task information and aid manipulation.…

Robotics · Computer Science 2018-10-16 Janis Stolzenwald , Walterio W. Mayol-Cuevas

Foundation models pre-trained on web-scale data are shown to encapsulate extensive world knowledge beneficial for robotic manipulation in the form of task planning. However, the actual physical implementation of these plans often relies on…

Robotics · Computer Science 2024-03-14 Haoxu Huang , Fanqi Lin , Yingdong Hu , Shengjie Wang , Yang Gao

Generalization remains a fundamental challenge in robotic manipulation. To tackle this challenge, recent Vision-Language-Action (VLA) models build policies on top of Vision-Language Models (VLMs), seeking to transfer their open-world…