中文
相关论文

相关论文: Hand Image Understanding via Deep Multi-Task Learn…

200 篇论文

In practical applications, computer vision tasks often need to be addressed simultaneously. Multitask learning typically achieves this by jointly training a single deep neural network to learn shared representations, providing efficiency…

计算机视觉与模式识别 · 计算机科学 2025-05-26 Konstantinos Spathis , Nikolaos Kardaris , Petros Maragos

Monocular 3D reconstruction of deformable objects, such as human body parts, has been typically approached by predicting parameters of heavyweight linear models. In this paper, we demonstrate an alternative solution that is based on the…

计算机视觉与模式识别 · 计算机科学 2019-08-06 Dominik Kulon , Haoyang Wang , Riza Alp Güler , Michael Bronstein , Stefanos Zafeiriou

Deploying real-time spatial perception on edge devices requires efficient multi-task models that leverage complementary task information while minimizing computational overhead. This paper introduces Multi-Mono-Hydra (M2H), a novel…

计算机视觉与模式识别 · 计算机科学 2026-05-19 U. V. B. L Udugama , George Vosselman , Francesco Nex

The segmentation-free research efforts for addressing handwritten text recognition can be divided into three categories: connectionist temporal classification (CTC), hidden Markov model and encoder-decoder methods. In this paper, inspired…

人工智能 · 计算机科学 2025-08-05 Zi-Rui Wang

While many individual tasks in the domain of human analysis have recently received an accuracy boost from deep learning approaches, multi-task learning has mostly been ignored due to a lack of data. New synthetic datasets are being…

计算机视觉与模式识别 · 计算机科学 2019-05-09 Daniel Sánchez , Marc Oliu , Meysam Madadi , Xavier Baró , Sergio Escalera

Text-to-image generation models have achieved remarkable advancements in recent years, aiming to produce realistic images from textual descriptions. However, these models often struggle with generating anatomically accurate representations…

计算机视觉与模式识别 · 计算机科学 2024-12-24 Haozhuo Zhang , Bin Zhu , Yu Cao , Yanbin Hao

The human hand moves in complex and high-dimensional ways, making estimation of 3D hand pose configurations from images alone a challenging task. In this work we propose a method to learn a statistical hand model represented by a…

计算机视觉与模式识别 · 计算机科学 2018-04-02 Adrian Spurr , Jie Song , Seonwook Park , Otmar Hilliges

Recent advancements in 3D hand pose estimation have shown promising results, but its effectiveness has primarily relied on the availability of large-scale annotated datasets, the creation of which is a laborious and costly process. To…

计算机视觉与模式识别 · 计算机科学 2023-08-16 Xiaozheng Zheng , Chao Wen , Zhou Xue , Pengfei Ren , Jingyu Wang

Hand pose estimation (HPE) can be used for a variety of human-computer interaction applications such as gesture-based control for physical or virtual/augmented reality devices. Recent works have shown that videos or multi-view images carry…

计算机视觉与模式识别 · 计算机科学 2021-09-27 Leyla Khaleghi , Alireza Sepas Moghaddam , Joshua Marshall , Ali Etemad

Human-Object Interaction (HOI) detection is a challenging computer vision task that requires visual models to address the complex interactive relationship between humans and objects and predict HOI triplets. Despite the challenges posed by…

计算机视觉与模式识别 · 计算机科学 2023-09-19 Yichao Cao , Qingfei Tang , Feng Yang , Xiu Su , Shan You , Xiaobo Lu , Chang Xu

Despite recent significant advancements in Handwritten Document Recognition (HDR), the efficient and accurate recognition of text against complex backgrounds, diverse handwriting styles, and varying document layouts remains a practical…

计算机视觉与模式识别 · 计算机科学 2025-01-23 Wenhao Gu , Li Gu , Ziqiang Wang , Ching Yee Suen , Yang Wang

Monocular 3D hand mesh recovery is challenging due to high degrees of freedom of hands, 2D-to-3D ambiguity and self-occlusion. Most existing methods are either inefficient or less straightforward for predicting the position of 3D mesh…

计算机视觉与模式识别 · 计算机科学 2025-12-08 Yihong Lin , Xianjia Wu , Xilai Wang , Jianqiao Hu , Songju Lei , Xiandong Li , Wenxiong Kang

Deep learning provides a new avenue for image restoration, which demands a delicate balance between fine-grained details and high-level contextualized information during recovering the latent clear image. In practice, however, existing…

计算机视觉与模式识别 · 计算机科学 2021-09-09 Man Zhou , Zeyu Xiao , Xueyang Fu , Aiping Liu , Gang Yang , Zhiwei Xiong

Visible and infrared image fusion (VIF) has attracted significant attention in recent years. Traditional VIF methods primarily focus on generating fused images with high visual quality, while recent advancements increasingly emphasize…

计算机视觉与模式识别 · 计算机科学 2025-05-13 Zixian Zhao , Andrew Howes , Xingchen Zhang

In general, hand pose estimation aims to improve the robustness of model performance in the real-world scenes. However, it is difficult to enhance the robustness since existing datasets are obtained in restricted environments to annotate 3D…

计算机视觉与模式识别 · 计算机科学 2023-10-20 Bosang Kim , Jonghyun Kim , Hyotae Lee , Lanying Jin , Jeongwon Ha , Dowoo Kwon , Jungpyo Kim , Wonhyeok Im , KyungMin Jin , Jungho Lee

Reconstructing a 3D hand from a single-view RGB image is challenging due to various hand configurations and depth ambiguity. To reliably reconstruct a 3D hand from a monocular image, most state-of-the-art methods heavily rely on 3D…

计算机视觉与模式识别 · 计算机科学 2021-03-23 Yujin Chen , Zhigang Tu , Di Kang , Linchao Bao , Ying Zhang , Xuefei Zhe , Ruizhi Chen , Junsong Yuan

Thanks to breakthroughs in AI and Deep learning methodology, Computer vision techniques are rapidly improving. Most computer vision applications require sophisticated image segmentation to comprehend what is image and to make an analysis of…

机器学习 · 计算机科学 2023-02-07 Lichun Gao , Chinmaya Khamesra , Uday Kumbhar , Ashay Aglawe

The capability to process multiple images is crucial for Large Vision-Language Models (LVLMs) to develop a more thorough and nuanced understanding of a scene. Recent multi-image LVLMs have begun to address this need. However, their…

计算机视觉与模式识别 · 计算机科学 2024-08-07 Fanqing Meng , Jin Wang , Chuanhao Li , Quanfeng Lu , Hao Tian , Jiaqi Liao , Xizhou Zhu , Jifeng Dai , Yu Qiao , Ping Luo , Kaipeng Zhang , Wenqi Shao

Reconstructing textured 3D human models from a single image is fundamental for AR/VR and digital human applications. However, existing methods mostly focus on single individuals and thus fail in multi-human scenes, where naive composition…

计算机视觉与模式识别 · 计算机科学 2026-04-08 Gwanghyun Kim , Junghun James Kim , Suh Yoon Jeon , Jason Park , Se Young Chun

Hyperspectral imaging acquires data in both the spatial and frequency domains to offer abundant physical or biological information. However, conventional hyperspectral imaging has intrinsic limitations of bulky instruments, slow data…

图像与视频处理 · 电气工程与系统科学 2023-04-06 Yuhyun Ji , Sang Mok Park , Semin Kwon , Jung Woo Leem , Vidhya Vijayakrishnan Nair , Yunjie Tong , Young L. Kim