English
Related papers

Related papers: Hourglass Tokenizer for Efficient Transformer-Base…

200 papers

We present Better Together, a method that simultaneously solves the human pose estimation problem while reconstructing a photorealistic 3D human avatar from multi-view videos. While prior art usually solves these problems separately, we…

Computer Vision and Pattern Recognition · Computer Science 2025-03-13 Arthur Moreau , Mohammed Brahimi , Richard Shaw , Athanasios Papaioannou , Thomas Tanay , Zhensong Zhang , Eduardo Pérez-Pellitero

Transformer encoder architectures have recently achieved state-of-the-art results on monocular 3D human mesh reconstruction, but they require a substantial number of parameters and expensive computations. Due to the large memory overhead…

Computer Vision and Pattern Recognition · Computer Science 2022-07-29 Junhyeong Cho , Kim Youwang , Tae-Hyun Oh

Action recognition and human pose estimation are closely related but both problems are generally handled as distinct tasks in the literature. In this work, we propose a multitask framework for jointly 2D and 3D pose estimation from still…

Computer Vision and Pattern Recognition · Computer Science 2018-03-22 Diogo C. Luvizon , David Picard , Hedi Tabia

Vision Transformers (ViT) have achieved remarkable success in large-scale image recognition. They split every 2D image into a fixed number of patches, each of which is treated as a token. Generally, representing an image with more tokens…

Computer Vision and Pattern Recognition · Computer Science 2021-10-27 Yulin Wang , Rui Huang , Shiji Song , Zeyi Huang , Gao Huang

While Transformers have rapidly gained popularity in various computer vision applications, post-hoc explanations of their internal mechanisms remain largely unexplored. Vision Transformers extract visual information by representing image…

Computer Vision and Pattern Recognition · Computer Science 2024-03-22 Junyi Wu , Bin Duan , Weitai Kang , Hao Tang , Yan Yan

Pose estimation commonly refers to computer vision methods that recognize people's body postures in images or videos. With recent advancements in deep learning, we now have compelling models to tackle the problem in real-time. Since these…

Robotics · Computer Science 2021-07-07 Arash Amini , Hafez Farazi , Sven Behnke

Traditional methods for human localization and pose estimation (HPE), which mainly rely on RGB images as an input modality, confront substantial limitations in real-world applications due to privacy concerns. In contrast, radar-based HPE…

Computer Vision and Pattern Recognition · Computer Science 2024-07-22 Yuan-Hao Ho , Jen-Hao Cheng , Sheng Yao Kuan , Zhongyu Jiang , Wenhao Chai , Hsiang-Wei Huang , Chih-Lung Lin , Jenq-Neng Hwang

Existing volumetric methods for predicting 3D human pose estimation are accurate, but computationally expensive and optimized for single time-step prediction. We present TEMPO, an efficient multi-view pose estimation model that learns a…

Computer Vision and Pattern Recognition · Computer Science 2023-09-15 Rohan Choudhury , Kris Kitani , Laszlo A. Jeni

Video-based human pose estimation models aim to address scenarios that cannot be effectively solved by static image models such as motion blur, out-of-focus and occlusion. Most existing approaches consist of two stages: detecting human…

Computer Vision and Pattern Recognition · Computer Science 2025-09-03 Zhihong Wei

The extraction of keypoint positions from input hand frames, known as 3D hand pose estimation, is crucial for various human-computer interaction applications. However, current approaches often struggle with the dynamic nature of…

Computer Vision and Pattern Recognition · Computer Science 2024-07-31 Wencan Cheng , Eunji Kim , Jong Hwan Ko

Real-time 6D object pose estimation is essential for many real-world applications, such as robotic grasping and augmented reality. To achieve an accurate object pose estimation from RGB images in real-time, we propose an effective and…

Computer Vision and Pattern Recognition · Computer Science 2022-04-21 Qi Guan , Zihao Sheng , Shibei Xue

Point tracking aims to localize corresponding points across video frames, serving as a fundamental task for 4D reconstruction, robotics, and video editing. Existing methods commonly rely on shallow convolutional backbones such as ResNet…

Computer Vision and Pattern Recognition · Computer Science 2025-12-24 Soowon Son , Honggyu An , Chaehyun Kim , Hyunah Ko , Jisu Nam , Dahyun Chung , Siyoon Jin , Jung Yi , Jaewon Min , Junhwa Hur , Seungryong Kim

3D hand pose estimation and shape recovery are challenging tasks in computer vision. We introduce a novel framework HandTailor, which combines a learning-based hand module and an optimization-based tailor module to achieve high-precision…

Computer Vision and Pattern Recognition · Computer Science 2021-10-25 Jun Lv , Wenqiang Xu , Lixin Yang , Sucheng Qian , Chongzhao Mao , Cewu Lu

Online video understanding is essential for applications like public surveillance and AI glasses. However, applying Multimodal Large Language Models (MLLMs) to this domain is challenging due to the large number of video frames, resulting in…

Computer Vision and Pattern Recognition · Computer Science 2025-12-16 Xinqi Jin , Hanxun Yu , Bohan Yu , Kebin Liu , Jian Liu , Keda Tao , Yixuan Pei , Huan Wang , Fan Dang , Jiangchuan Liu , Weiqiang Wang

Human-Object Interaction (HOI) detection is a task of identifying "a set of interactions" in an image, which involves the i) localization of the subject (i.e., humans) and target (i.e., objects) of interaction, and ii) the classification of…

Computer Vision and Pattern Recognition · Computer Science 2021-04-29 Bumsoo Kim , Junhyun Lee , Jaewoo Kang , Eun-Sol Kim , Hyunwoo J. Kim

Learning based 6D object pose estimation methods rely on computing large intermediate pose representations and/or iteratively refining an initial estimation with a slow render-compare pipeline. This paper introduces a novel method we call…

Computer Vision and Pattern Recognition · Computer Science 2022-10-24 Pedro Castro , Tae-Kyun Kim

This paper presents a new vision Transformer, named Iwin Transformer, which is specifically designed for human-object interaction (HOI) detection, a detailed scene understanding task involving a sequential process of human/object detection…

Computer Vision and Pattern Recognition · Computer Science 2022-10-21 Danyang Tu , Xiongkuo Min , Huiyu Duan , Guodong Guo , Guangtao Zhai , Wei Shen

This study presents a new network (i.e., PoseLifter) that can lift a 2D human pose to an absolute 3D pose in a camera coordinate system. The proposed network estimates the absolute 3D location of a target subject and generates an improved…

Computer Vision and Pattern Recognition · Computer Science 2020-03-16 Ju Yong Chang , Gyeongsik Moon , Kyoung Mu Lee

3D Human body pose and shape estimation within a temporal sequence can be quite critical for understanding human behavior. Despite the significant progress in human pose estimation in the recent years, which are often based on single images…

Computer Vision and Pattern Recognition · Computer Science 2022-07-27 Zhouping Wang , Sarah Ostadabbas

Token pruning is essential for enhancing the computational efficiency of vision-language models (VLMs), particularly for video-based tasks where temporal redundancy is prevalent. Prior approaches typically prune tokens either (1) within the…

Computer Vision and Pattern Recognition · Computer Science 2026-03-19 Jianrui Zhang , Yue Yang , Rohun Tripathi , Winson Han , Ranjay Krishna , Christopher Clark , Yong Jae Lee , Sangho Lee
‹ Prev 1 8 9 10 Next ›