English
Related papers

Related papers: MaskHOI: Robust 3D Hand-Object Interaction Estimat…

200 papers

Recently, 3D hand reconstruction has gained more attention in human-computer cooperation, especially for hand-object interaction scenario. However, it still remains huge challenge due to severe hand-occlusion caused by interaction, which…

Computer Vision and Pattern Recognition · Computer Science 2024-03-05 Feng Shuang , Wenbo He , Shaodong Li

Existing methods for 3D tracking from monocular RGB videos predominantly consider articulated and rigid objects. Modelling dense non-rigid object deformations in this setting remained largely unaddressed so far, although such effects can…

Computer Vision and Pattern Recognition · Computer Science 2023-10-16 Soshi Shimada , Vladislav Golyanik , Patrick Pérez , Christian Theobalt

Many manipulation tasks, such as placement or within-hand manipulation, require the object's pose relative to a robot hand. The task is difficult when the hand significantly occludes the object. It is especially hard for adaptive hands, for…

Robotics · Computer Science 2021-12-20 Bowen Wen , Chaitanya Mitash , Sruthi Soorian , Andrew Kimmel , Avishai Sintov , Kostas E. Bekris

This paper shows that masked autoencoders (MAE) are scalable self-supervised learners for computer vision. Our MAE approach is simple: we mask random patches of the input image and reconstruct the missing pixels. It is based on two core…

Computer Vision and Pattern Recognition · Computer Science 2021-12-21 Kaiming He , Xinlei Chen , Saining Xie , Yanghao Li , Piotr Dollár , Ross Girshick

We ask whether everyday open-world monocular videos can be turned into reusable 4D interaction primitives: articulated hand motion, object shape with 6D pose over time, and the when/where of contact. Such a capability would enable scalable…

Computer Vision and Pattern Recognition · Computer Science 2026-05-22 Hao Xu , Yilin Liu , Yinqiao Wang , Chi-Wing Fu , Niloy J. Mitra

Recent general-purpose audio representations show state-of-the-art performance on various audio tasks. These representations are pre-trained by self-supervised learning methods that create training signals from the input. For example,…

Audio and Speech Processing · Electrical Eng. & Systems 2023-03-09 Daisuke Niizumi , Daiki Takeuchi , Yasunori Ohishi , Noboru Harada , Kunio Kashino

Masked autoencoders (MAE) have recently been introduced to 3D self-supervised pretraining for point clouds due to their great success in NLP and computer vision. Unlike MAEs used in the image domain, where the pretext task is to restore…

Computer Vision and Pattern Recognition · Computer Science 2024-04-30 Siming Yan , Yuqi Yang , Yuxiao Guo , Hao Pan , Peng-shuai Wang , Xin Tong , Yang Liu , Qixing Huang

Hand pose estimation is a fundamental task in many human-robot interaction-related applications. However, previous approaches suffer from unsatisfying hand landmark predictions in real-world scenes and high computation burden. This paper…

Robotics · Computer Science 2021-10-13 Shan An , Xiajie Zhang , Dong Wei , Haogang Zhu , Jianyu Yang , Konstantinos A. Tsintotas

Understanding and recognizing human-object interaction (HOI) is a pivotal application in AR/VR and robotics. Recent open-vocabulary HOI detection approaches depend exclusively on large language models for richer textual prompts, neglecting…

Computer Vision and Pattern Recognition · Computer Science 2025-08-18 Zhenhao Zhang , Hanqing Wang , Xiangyu Zeng , Ziyu Cheng , Jiaxin Liu , Haoyu Yan , Zhirui Liu , Kaiyang Ji , Tianxiang Gui , Ke Hu , Kangyi Chen , Yahao Fan , Mokai Pan

In this paper, we develop \textbf{MP-HOI}, a powerful Multi-modal Prompt-based HOI detector designed to leverage both textual descriptions for open-set generalization and visual exemplars for handling high ambiguity in descriptions,…

Computer Vision and Pattern Recognition · Computer Science 2024-06-12 Jie Yang , Bingliang Li , Ailing Zeng , Lei Zhang , Ruimao Zhang

We present a unified framework for understanding 3D hand and object interactions in raw image sequences from egocentric RGB cameras. Given a single RGB image, our model jointly estimates the 3D hand and object poses, models their…

Computer Vision and Pattern Recognition · Computer Science 2019-04-11 Bugra Tekin , Federica Bogo , Marc Pollefeys

Hand manipulating objects is an important interaction motion in our daily activities. We faithfully reconstruct this motion with a single RGBD camera by a novel deep reinforcement learning method to leverage physics. Firstly, we propose…

Computer Vision and Pattern Recognition · Computer Science 2024-05-07 Haoyu Hu , Xinyu Yi , Zhe Cao , Jun-Hai Yong , Feng Xu

The sensing process of large-scale LiDAR point clouds inevitably causes large blind spots, i.e. regions not visible to the sensor. We demonstrate how these inherent sampling properties can be effectively utilized for self-supervised…

Computer Vision and Pattern Recognition · Computer Science 2023-12-08 Georg Krispel , David Schinagl , Christian Fruhwirth-Reisinger , Horst Possegger , Horst Bischof

Humans inherently possess generalizable visual representations that empower them to efficiently explore and interact with the environments in manipulation tasks. We advocate that such a representation automatically arises from…

Fully supervised skeleton-based action recognition has achieved great progress with the blooming of deep learning techniques. However, these methods require sufficient labeled data which is not easy to obtain. In contrast, self-supervised…

Computer Vision and Pattern Recognition · Computer Science 2023-05-12 Wenhan Wu , Yilei Hua , Ce Zheng , Shiqian Wu , Chen Chen , Aidong Lu

We present a new pre-training strategy called M$^{3}$3D ($\underline{M}$ulti-$\underline{M}$odal $\underline{M}$asked $\underline{3D}$) built based on Multi-modal masked autoencoders that can leverage 3D priors and learned cross-modal…

Computer Vision and Pattern Recognition · Computer Science 2023-09-28 Muhammad Abdullah Jamal , Omid Mohareri

Hands are often severely occluded by objects, which makes 3D hand mesh estimation challenging. Previous works often have disregarded information at occluded regions. However, we argue that occluded regions have strong correlations with…

Computer Vision and Pattern Recognition · Computer Science 2022-03-29 JoonKyu Park , Yeonguk Oh , Gyeongsik Moon , Hongsuk Choi , Kyoung Mu Lee

Modeling human-object interactions (HOI) from an egocentric perspective is a critical yet challenging task, particularly when relying on sparse signals from wearable devices like smart glasses and watches. We present ECHO, the first unified…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Ilya A. Petrov , Vladimir Guzov , Riccardo Marin , Emre Aksan , Xu Chen , Daniel Cremers , Thabo Beeler , Gerard Pons-Moll

Recently, various studies have been directed towards exploring dense passage retrieval techniques employing pre-trained language models, among which the masked auto-encoder (MAE) pre-training architecture has emerged as the most promising.…

Information Retrieval · Computer Science 2023-05-23 Zehan Li , Yanzhao Zhang , Dingkun Long , Pengjun Xie

Articulated hand pose estimation is a challenging task for human-computer interaction. The state-of-the-art hand pose estimation algorithms work only with one or a few subjects for which they have been calibrated or trained. Particularly,…

Human-Computer Interaction · Computer Science 2017-12-11 Jameel Malik , Ahmed Elhayek , Didier Stricker