English
Related papers

Related papers: H2O: A Benchmark for Visual Human-human Object Han…

200 papers

Much of the literature on robotic perception focuses on the visual modality. Vision provides a global observation of a scene, making it broadly useful. However, in the domain of robotic manipulation, vision alone can sometimes prove…

Robotics · Computer Science 2019-03-11 Justin Lin , Roberto Calandra , Sergey Levine

We present an approach to learn general robot manipulation priors from 3D hand-object interaction trajectories. We build a framework to use in-the-wild videos to generate sensorimotor robot trajectories. We do so by lifting both the human…

Jointly estimating hand and object shape facilitates the grasping task in human-to-robot handovers. However, relying on hand-crafted prior knowledge about the geometric structure of the object fails when generalising to unseen objects, and…

Robotics · Computer Science 2025-05-13 Yik Lung Pang , Alessio Xompero , Changjae Oh , Andrea Cavallaro

Knowledge of the 6D pose of an object can benefit in-hand object manipulation. In-hand 6D object pose estimation is challenging because of heavy occlusion produced by the robot's grippers, which can have an adverse effect on methods that…

We introduce the Lecture Video Visual Objects (LVVO) dataset, a new benchmark for visual object detection in educational video content. The dataset consists of 4,000 frames extracted from 245 lecture videos spanning biology, computer…

Computer Vision and Pattern Recognition · Computer Science 2025-06-18 Dipayan Biswas , Shishir Shah , Jaspal Subhlok

Due to real-world dynamics and hardware uncertainty, robots inevitably fail in task executions, resulting in undesired or even dangerous executions. In order to avoid failures and improve robot performance, it is critical to identify and…

Robotics · Computer Science 2021-06-30 Boyi Song , Yuntao Peng , Ruijiao Luo , Rui Liu

Grasping objects with limited or no prior knowledge about them is a highly relevant skill in assistive robotics. Still, in this general setting, it has remained an open problem, especially when it comes to only partial observability and…

Robotics · Computer Science 2026-01-21 Matthias Humt , Dominik Winkelbauer , Ulrich Hillenbrand , Berthold Bäuml

Can a robot grasp an unknown object without seeing it? In this paper, we present a tactile-sensing based approach to this challenging problem of grasping novel objects without prior knowledge of their location or physical properties. Our…

Robotics · Computer Science 2018-05-14 Adithyavairavan Murali , Yin Li , Dhiraj Gandhi , Abhinav Gupta

This paper proposes 12 multi-object grasps (MOGs) types from a human and robot grasping data set. The grasp types are then analyzed and organized into a MOG taxonomy. This paper first presents three MOG data collection setups: a human…

Robotics · Computer Science 2022-05-31 Yu Sun , Eliza Amatova , Tianze Chen

Multimodal object recognition is still an emerging field. Thus, publicly available datasets are still rare and of small size. This dataset was developed to help fill this void and presents multimodal data for 63 objects with some visual and…

Implicit communication plays such a crucial role during social exchanges that it must be considered for a good experience in human-robot interaction. This work addresses implicit communication associated with the detection of physical…

We describe a learning-based approach to hand-eye coordination for robotic grasping from monocular images. To learn hand-eye coordination for grasping, we trained a large convolutional neural network to predict the probability that…

Machine Learning · Computer Science 2016-08-30 Sergey Levine , Peter Pastor , Alex Krizhevsky , Deirdre Quillen

Supernumerary robotic limbs (SRLs) are robotic structures integrated closely with the user's body, which augment human physical capabilities and necessitate seamless, naturalistic human-machine interaction. For effective assistance in…

The evaluation of object detection models is usually performed by optimizing a single metric, e.g. mAP, on a fixed set of datasets, e.g. Microsoft COCO and Pascal VOC. Due to image retrieval and annotation costs, these datasets consist…

Computer Vision and Pattern Recognition · Computer Science 2022-12-01 Floriana Ciaglia , Francesco Saverio Zuppichini , Paul Guerrie , Mark McQuade , Jacob Solawetz

This paper addresses the task of detecting and recognizing human-object interactions (HOI) in images and videos. We introduce the Graph Parsing Neural Network (GPNN), a framework that incorporates structural knowledge while being…

Computer Vision and Pattern Recognition · Computer Science 2018-08-27 Siyuan Qi , Wenguan Wang , Baoxiong Jia , Jianbing Shen , Song-Chun Zhu

We address key limitations in existing datasets and models for task-oriented hand-object interaction video generation, a critical approach of generating video demonstrations for robotic imitation learning. Current datasets, such as Ego4D,…

Computer Vision and Pattern Recognition · Computer Science 2025-06-09 Hongxiang Zhao , Xingchen Liu , Mutian Xu , Yiming Hao , Weikai Chen , Xiaoguang Han

Comprehensive visual understanding requires detection frameworks that can effectively learn and utilize object interactions while analyzing objects individually. This is the main objective in Human-Object Interaction (HOI) detection task.…

Computer Vision and Pattern Recognition · Computer Science 2020-03-13 Oytun Ulutan , A S M Iftekhar , B. S. Manjunath

Human object interaction (HOI) detection is an important task in image understanding and reasoning. It is in a form of HOI triplet <human; verb; object>, requiring bounding boxes for human and object, and action between them for the task…

Computer Vision and Pattern Recognition · Computer Science 2020-11-13 Suresh Kirthi Kumaraswamy , Miaojing Shi , Ewa Kijak

Human-robot collaboration requires the contactless estimation of the physical properties of containers manipulated by a person, for example while pouring content in a cup or moving a food box. Acoustic and visual signals can be used to…

Multimedia · Computer Science 2022-03-07 A. Xompero , Y. L. Pang , T. Patten , A. Prabhakar , B. Calli , A. Cavallaro

This paper proposes an interaction reasoning network for modelling spatio-temporal relationships between hands and objects in video. The proposed interaction unit utilises a Transformer module to reason about each acting hand, and its…

Computer Vision and Pattern Recognition · Computer Science 2022-01-14 Jian Ma , Dima Damen