English
Related papers

Related papers: TACO: Benchmarking Generalizable Bimanual Tool-ACt…

200 papers

Modeling human-object interactions (HOI) from an egocentric perspective is a critical yet challenging task, particularly when relying on sparse signals from wearable devices like smart glasses and watches. We present ECHO, the first unified…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Ilya A. Petrov , Vladimir Guzov , Riccardo Marin , Emre Aksan , Xu Chen , Daniel Cremers , Thabo Beeler , Gerard Pons-Moll

When your robot grasps an object using dexterous hands or grippers, it should understand the Task-Oriented Affordances of the Object(TOAO), as different tasks often require attention to specific parts of the object. To address this…

Robotics · Computer Science 2024-09-19 Jiawen Wang , Dingsheng Luo

Many advanced Learning from Demonstration (LfD) methods consider the decomposition of complex, real-world tasks into simpler sub-tasks. By reusing the corresponding sub-policies within and between tasks, they provide training data for each…

Machine Learning · Computer Science 2018-08-13 Kyriacos Shiarlis , Markus Wulfmeier , Sasha Salter , Shimon Whiteson , Ingmar Posner

We present a novel approach for hand-object action recognition that leverages 2D point tracks as an additional motion cue. While most existing methods rely on RGB appearance, human pose estimation, or their combination, our work…

Computer Vision and Pattern Recognition · Computer Science 2026-01-12 Dennis Holzmann , Sven Wachsmuth

Integrating robots into human-centric environments such as homes, necessitates advanced manipulation skills as robotic devices will need to engage with articulated objects like doors and drawers. Key challenges in robotic manipulation of…

Robotics · Computer Science 2025-09-15 Zihang Zhao , Yuyang Li , Wanlin Li , Zhenghao Qi , Lecheng Ruan , Yixin Zhu , Kaspar Althoefer

Digital interaction with everyday objects has become popular since the proliferation of camera-based systems that detect and augment objects "just-in-time". Common systems use a vision-based approach to detect objects and display their…

Human-Computer Interaction · Computer Science 2020-12-22 Thomas Kosch , Albrecht Schmidt

Human video demonstrations provide abundant training data for learning robot policies, but video alone cannot capture the rich contact signals critical for mastering manipulation. We introduce OSMO, an open-source wearable tactile glove…

We present a system for learning generalizable hand-object tracking controllers purely from synthetic data, without requiring any human demonstrations. Our approach makes two key contributions: (1) HOP, a Hand-Object Planner, which can…

Robotics · Computer Science 2025-12-23 Yinhuai Wang , Runyi Yu , Hok Wai Tsui , Xiaoyi Lin , Hui Zhang , Qihan Zhao , Ke Fan , Miao Li , Jie Song , Jingbo Wang , Qifeng Chen , Ping Tan

Estimating 3D hand and object pose from a single image is an extremely challenging problem: hands and objects are often self-occluded during interactions, and the 3D annotations are scarce as even humans cannot directly label the…

Computer Vision and Pattern Recognition · Computer Science 2021-06-10 Shaowei Liu , Hanwen Jiang , Jiarui Xu , Sifei Liu , Xiaolong Wang

Spatio-temporal Human-Object Interaction (ST-HOI) understanding aims at detecting HOIs from videos, which is crucial for activity understanding. However, existing whole-body-object interaction video benchmarks overlook the truth that…

Computer Vision and Pattern Recognition · Computer Science 2025-02-25 Xiaoyang Liu , Boran Wen , Xinpeng Liu , Zizheng Zhou , Hongwei Fan , Cewu Lu , Lizhuang Ma , Yulong Chen , Yong-Lu Li

Vision-Language-Action (VLA) models have recently emerged as powerful generalists for robotic manipulation. However, due to their predominant reliance on visual modalities, they fundamentally lack the physical intuition required for…

Robotics · Computer Science 2026-02-02 Yuzhe Huang , Pei Lin , Wanlin Li , Daohan Li , Jiajun Li , Jiaming Jiang , Chenxi Xiao , Ziyuan Jiao

Activity recognition has shown impressive progress in recent years. However, the challenges of detecting fine-grained activities and understanding how they are combined into composite activities have been largely overlooked. In this work we…

Computer Vision and Pattern Recognition · Computer Science 2015-10-16 Marcus Rohrbach , Anna Rohrbach , Michaela Regneri , Sikandar Amin , Mykhaylo Andriluka , Manfred Pinkal , Bernt Schiele

Bimanual human activities inherently involve coordinated movements of both hands and body. However, the impact of this coordination in activity understanding has not been systematically evaluated due to the lack of suitable datasets. Such…

Computer Vision and Pattern Recognition · Computer Science 2025-09-30 Tatsuro Banno , Takehiko Ohkawa , Ruicong Liu , Ryosuke Furuta , Yoichi Sato

Tactile and visual perception are both crucial for humans to perform fine-grained interactions with their environment. Developing similar multi-modal sensing capabilities for robots can significantly enhance and expand their manipulation…

Robotics · Computer Science 2025-01-08 Binghao Huang , Yixuan Wang , Xinyi Yang , Yiyue Luo , Yunzhu Li

Amodal perception, the ability to comprehend complete object structures from partial visibility, is a fundamental skill, even for infants. Its significance extends to applications like autonomous driving, where a clear understanding of…

Computer Vision and Pattern Recognition · Computer Science 2024-04-04 Cheng-Yen Hsieh , Kaihua Chen , Achal Dave , Tarasha Khurana , Deva Ramanan

Egocentric hand-object motion generation is crucial for immersive AR/VR and robotic imitation but remains challenging due to unstable viewpoints, self-occlusions, perspective distortion, and noisy ego-motion. Existing methods rely on…

Computer Vision and Pattern Recognition · Computer Science 2025-05-23 Bohan Zhou , Yi Zhan , Zhongbin Zhang , Zongqing Lu

Human-Object Interaction (HOI) detection aims to identify humans and objects within images and interpret their interactions. Existing HOI methods rely heavily on large datasets with manual annotations to learn interactions from visual cues.…

Computer Vision and Pattern Recognition · Computer Science 2025-07-24 Francesco Tonini , Lorenzo Vaquero , Alessandro Conti , Cigdem Beyan , Elisa Ricci

Object models are gradually progressing from predicting just category labels to providing detailed descriptions of object instances. This motivates the need for large datasets which go beyond traditional object masks and provide richer…

Tracking and detecting any object, including ones never-seen-before during model training, is a crucial but elusive capability of autonomous systems. An autonomous agent that is blind to never-seen-before objects poses a safety hazard when…

Computer Vision and Pattern Recognition · Computer Science 2022-03-30 Yang Liu , Idil Esen Zulfikar , Jonathon Luiten , Achal Dave , Deva Ramanan , Bastian Leibe , Aljoša Ošep , Laura Leal-Taixé

Bimanual robotic manipulation is an emerging and critical topic in the robotics community. Previous works primarily rely on integrated control models that take the perceptions and states of both arms as inputs to directly predict their…

Robotics · Computer Science 2025-11-05 Jian-Jian Jiang , Xiao-Ming Wu , Yi-Xiang He , Ling-An Zeng , Yi-Lin Wei , Dandan Zhang , Wei-Shi Zheng