English
Related papers

Related papers: Optimizing Hand Region Detection in MediaPipe Holi…

200 papers

The success of Deep Convolutional Neural Networks (CNNs) in recent years in almost all the Computer Vision tasks on one hand, and the popularity of low-cost consumer depth cameras on the other, has made Hand Pose Estimation a hot topic in…

Computer Vision and Pattern Recognition · Computer Science 2019-06-04 Bardia Doosti

Dexterous manipulation through imitation learning has gained significant attention in robotics research. The collection of high-quality expert data holds paramount importance when using imitation learning. The existing approaches for…

Robotics · Computer Science 2023-09-27 Dehao Wei , Huazhe Xu

With the increase number of companies focusing on commercializing Augmented Reality (AR), Virtual Reality (VR) and wearable devices, the need for a hand based input mechanism is becoming essential in order to make the experience natural,…

Computer Vision and Pattern Recognition · Computer Science 2016-04-22 Emad Barsoum

Contrastive Language-Image Pre-training (CLIP) has demonstrated remarkable generalization ability and strong performance across a wide range of vision-language tasks. However, due to the lack of region-level supervision, CLIP exhibits…

Computer Vision and Pattern Recognition · Computer Science 2025-12-01 Haoxi Zeng , Haoxuan Li , Yi Bin , Pengpeng Zeng , Xing Xu , Yang Yang , Heng Tao Shen

In grasp detection, the robot estimates the position and orientation of potential grasp configurations directly from sensor data. This paper explores the relationship between viewpoint and grasp detection performance. Specifically, we…

Robotics · Computer Science 2017-08-01 Marcus Gualtieri , Robert Platt

Dexterous robotic hands require high-speed multimodal sensing across many degrees of freedom, yet existing readout architectures often impose trade-offs between sensor count, wiring complexity, and sampling bandwidth. This paper presents a…

Sign language recognition (SLR) systems typically require large labeled corpora for each language, yet the majority of the world's 300+ sign languages lack sufficient annotated data. Cross-lingual few-shot transfer, pretraining on a…

Computer Vision and Pattern Recognition · Computer Science 2026-03-11 Chayanin Chamachot , Kanokphan Lertniponphan

We propose a camera-based assistive text reading framework to help blind persons read text labels and product packaging from hand-held objects in their daily life. To isolate the object from untidy backgrounds or other surrounding objects…

Human-Computer Interaction · Computer Science 2019-01-18 Rajkumar N , Anand M. G , Barathiraja N

Robotic dexterous manipulation is a challenging problem due to high degrees of freedom (DoFs) and complex contacts of multi-fingered robotic hands. Many existing deep reinforcement learning (DRL) based methods aim at improving sample…

Robotics · Computer Science 2026-02-26 Qingtao Liu , Zhengnan Sun , Yu Cui , Haoming Li , Gaofeng Li , Lin Shao , Jiming Chen , Qi Ye

Hand pose estimation from 3D depth images, has been explored widely using various kinds of techniques in the field of computer vision. Though, deep learning based method improve the performance greatly recently, however, this problem still…

Computer Vision and Pattern Recognition · Computer Science 2020-01-24 Zhaohui Zhang , Shipeng Xie , Mingxiu Chen , Haichao Zhu

6D pose confidence region estimation has emerged as a critical direction, aiming to perform uncertainty quantification for assessing the reliability of estimated poses. However, current sampling-based approach suffers from critical…

Computer Vision and Pattern Recognition · Computer Science 2025-07-01 Jinghao Wang , Zhang Li , Zi Wang , Banglei Guan , Yang Shang , Qifeng Yu

Human object interaction (HOI) detection plays a crucial role in human-centric scene understanding and serves as a fundamental building-block for many vision tasks. One generalizable and scalable strategy for HOI detection is to use weak…

Computer Vision and Pattern Recognition · Computer Science 2023-03-03 Bo Wan , Yongfei Liu , Desen Zhou , Tinne Tuytelaars , Xuming He

Great progress has been made in learning-based object detection methods in the last decade. Two-stage detectors often have higher detection accuracy than one-stage detectors, due to the use of region of interest (RoI) feature extractors…

Computer Vision and Pattern Recognition · Computer Science 2023-12-18 Guo-Ye Yang , George Kiyohiro Nakayama , Zi-Kai Xiao , Tai-Jiang Mu , Xiaolei Huang , Shi-Min Hu

3D hand-object pose estimation is the key to the success of many computer vision applications. The main focus of this task is to effectively model the interaction between the hand and an object. To this end, existing works either rely on…

Computer Vision and Pattern Recognition · Computer Science 2023-01-09 Rong Wang , Wei Mao , Hongdong Li

We focus on the task of everyday hand pose estimation from egocentric viewpoints. For this task, we show that depth sensors are particularly informative for extracting near-field interactions of the camera wearer with his/her environment.…

Computer Vision and Pattern Recognition · Computer Science 2014-12-02 Gregory Rogez , James S. Supancic , Maryam Khademi , Jose Maria Martinez Montiel , Deva Ramanan

Regions of Interest (ROI) contain morphological features in pathology whole slide images (WSI) are delimited with polygons[1]. These polygons are often represented in either a textual notation (with the array of edges) or in a binary mask…

Graphics · Computer Science 2020-05-15 Erich Bremer , Jonas Almeida , Joel Saltz

Human hand actions are quite complex, especially when they involve object manipulation, mainly due to the high dimensionality of the hand and the vast action space that entails. Imitating those actions with dexterous hand models involves…

Computer Vision and Pattern Recognition · Computer Science 2018-10-04 Dafni Antotsiou , Guillermo Garcia-Hernando , Tae-Kyun Kim

A large number of works in egocentric vision have concentrated on action and object recognition. Detection and segmentation of hands in first-person videos, however, has less been explored. For many applications in this domain, it is…

Computer Vision and Pattern Recognition · Computer Science 2018-03-30 Aisha Urooj Khan , Ali Borji

Transformers rely on explicit positional encoding to model structure in data. While Rotary Position Embedding (RoPE) excels in 1D domains, its application to image generation reveals significant limitations such as fine-grained spatial…

Computer Vision and Pattern Recognition · Computer Science 2026-03-13 Jiaye Li , Baoyou Chen , Hui Li , Zilong Dong , Jingdong Wang , Siyu Zhu

Estimating 3D hand pose from 2D images is a difficult, inverse problem due to the inherent scale and depth ambiguities. Current state-of-the-art methods train fully supervised deep neural networks with 3D ground-truth data. However,…

Computer Vision and Pattern Recognition · Computer Science 2020-08-05 Adrian Spurr , Umar Iqbal , Pavlo Molchanov , Otmar Hilliges , Jan Kautz