中文
相关论文

相关论文: Optimizing Hand Region Detection in MediaPipe Holi…

200 篇论文

The success of Deep Convolutional Neural Networks (CNNs) in recent years in almost all the Computer Vision tasks on one hand, and the popularity of low-cost consumer depth cameras on the other, has made Hand Pose Estimation a hot topic in…

计算机视觉与模式识别 · 计算机科学 2019-06-04 Bardia Doosti

Dexterous manipulation through imitation learning has gained significant attention in robotics research. The collection of high-quality expert data holds paramount importance when using imitation learning. The existing approaches for…

机器人学 · 计算机科学 2023-09-27 Dehao Wei , Huazhe Xu

With the increase number of companies focusing on commercializing Augmented Reality (AR), Virtual Reality (VR) and wearable devices, the need for a hand based input mechanism is becoming essential in order to make the experience natural,…

计算机视觉与模式识别 · 计算机科学 2016-04-22 Emad Barsoum

Contrastive Language-Image Pre-training (CLIP) has demonstrated remarkable generalization ability and strong performance across a wide range of vision-language tasks. However, due to the lack of region-level supervision, CLIP exhibits…

计算机视觉与模式识别 · 计算机科学 2025-12-01 Haoxi Zeng , Haoxuan Li , Yi Bin , Pengpeng Zeng , Xing Xu , Yang Yang , Heng Tao Shen

In grasp detection, the robot estimates the position and orientation of potential grasp configurations directly from sensor data. This paper explores the relationship between viewpoint and grasp detection performance. Specifically, we…

机器人学 · 计算机科学 2017-08-01 Marcus Gualtieri , Robert Platt

Dexterous robotic hands require high-speed multimodal sensing across many degrees of freedom, yet existing readout architectures often impose trade-offs between sensor count, wiring complexity, and sampling bandwidth. This paper presents a…

Sign language recognition (SLR) systems typically require large labeled corpora for each language, yet the majority of the world's 300+ sign languages lack sufficient annotated data. Cross-lingual few-shot transfer, pretraining on a…

计算机视觉与模式识别 · 计算机科学 2026-03-11 Chayanin Chamachot , Kanokphan Lertniponphan

We propose a camera-based assistive text reading framework to help blind persons read text labels and product packaging from hand-held objects in their daily life. To isolate the object from untidy backgrounds or other surrounding objects…

人机交互 · 计算机科学 2019-01-18 Rajkumar N , Anand M. G , Barathiraja N

Robotic dexterous manipulation is a challenging problem due to high degrees of freedom (DoFs) and complex contacts of multi-fingered robotic hands. Many existing deep reinforcement learning (DRL) based methods aim at improving sample…

机器人学 · 计算机科学 2026-02-26 Qingtao Liu , Zhengnan Sun , Yu Cui , Haoming Li , Gaofeng Li , Lin Shao , Jiming Chen , Qi Ye

Hand pose estimation from 3D depth images, has been explored widely using various kinds of techniques in the field of computer vision. Though, deep learning based method improve the performance greatly recently, however, this problem still…

计算机视觉与模式识别 · 计算机科学 2020-01-24 Zhaohui Zhang , Shipeng Xie , Mingxiu Chen , Haichao Zhu

6D pose confidence region estimation has emerged as a critical direction, aiming to perform uncertainty quantification for assessing the reliability of estimated poses. However, current sampling-based approach suffers from critical…

计算机视觉与模式识别 · 计算机科学 2025-07-01 Jinghao Wang , Zhang Li , Zi Wang , Banglei Guan , Yang Shang , Qifeng Yu

Human object interaction (HOI) detection plays a crucial role in human-centric scene understanding and serves as a fundamental building-block for many vision tasks. One generalizable and scalable strategy for HOI detection is to use weak…

计算机视觉与模式识别 · 计算机科学 2023-03-03 Bo Wan , Yongfei Liu , Desen Zhou , Tinne Tuytelaars , Xuming He

Great progress has been made in learning-based object detection methods in the last decade. Two-stage detectors often have higher detection accuracy than one-stage detectors, due to the use of region of interest (RoI) feature extractors…

计算机视觉与模式识别 · 计算机科学 2023-12-18 Guo-Ye Yang , George Kiyohiro Nakayama , Zi-Kai Xiao , Tai-Jiang Mu , Xiaolei Huang , Shi-Min Hu

3D hand-object pose estimation is the key to the success of many computer vision applications. The main focus of this task is to effectively model the interaction between the hand and an object. To this end, existing works either rely on…

计算机视觉与模式识别 · 计算机科学 2023-01-09 Rong Wang , Wei Mao , Hongdong Li

We focus on the task of everyday hand pose estimation from egocentric viewpoints. For this task, we show that depth sensors are particularly informative for extracting near-field interactions of the camera wearer with his/her environment.…

计算机视觉与模式识别 · 计算机科学 2014-12-02 Gregory Rogez , James S. Supancic , Maryam Khademi , Jose Maria Martinez Montiel , Deva Ramanan

Regions of Interest (ROI) contain morphological features in pathology whole slide images (WSI) are delimited with polygons[1]. These polygons are often represented in either a textual notation (with the array of edges) or in a binary mask…

图形学 · 计算机科学 2020-05-15 Erich Bremer , Jonas Almeida , Joel Saltz

Human hand actions are quite complex, especially when they involve object manipulation, mainly due to the high dimensionality of the hand and the vast action space that entails. Imitating those actions with dexterous hand models involves…

计算机视觉与模式识别 · 计算机科学 2018-10-04 Dafni Antotsiou , Guillermo Garcia-Hernando , Tae-Kyun Kim

A large number of works in egocentric vision have concentrated on action and object recognition. Detection and segmentation of hands in first-person videos, however, has less been explored. For many applications in this domain, it is…

计算机视觉与模式识别 · 计算机科学 2018-03-30 Aisha Urooj Khan , Ali Borji

Transformers rely on explicit positional encoding to model structure in data. While Rotary Position Embedding (RoPE) excels in 1D domains, its application to image generation reveals significant limitations such as fine-grained spatial…

计算机视觉与模式识别 · 计算机科学 2026-03-13 Jiaye Li , Baoyou Chen , Hui Li , Zilong Dong , Jingdong Wang , Siyu Zhu

Estimating 3D hand pose from 2D images is a difficult, inverse problem due to the inherent scale and depth ambiguities. Current state-of-the-art methods train fully supervised deep neural networks with 3D ground-truth data. However,…

计算机视觉与模式识别 · 计算机科学 2020-08-05 Adrian Spurr , Umar Iqbal , Pavlo Molchanov , Otmar Hilliges , Jan Kautz