English
Related papers

Related papers: EMO-X: Efficient Multi-Person Pose and Shape Estim…

200 papers

Annotating biomedical images for supervised learning is a complex and labor-intensive task due to data diversity and its intricate nature. In this paper, we propose an innovative method, the efficient one-pass selective annotation (EPOSA),…

Image and Video Processing · Electrical Eng. & Systems 2023-09-18 Yuli Wang , Peiyu Duan , Zhangxing Bian , Anqi Feng , Yuan Xue

People touch their face 23 times an hour, they cross their arms and legs, put their hands on their hips, etc. While many images of people contain some form of self-contact, current 3D human pose and shape (HPS) regression methods typically…

Computer Vision and Pattern Recognition · Computer Science 2021-04-09 Lea Müller , Ahmed A. A. Osman , Siyu Tang , Chun-Hao P. Huang , Michael J. Black

In this paper, we introduce MeshMamba, a neural network model for learning 3D articulated mesh models by employing the recently proposed Mamba State Space Models (Mamba-SSMs). MeshMamba is efficient and scalable in handling a large number…

Computer Vision and Pattern Recognition · Computer Science 2025-07-22 Yusuke Yoshiyasu , Leyuan Sun , Ryusuke Sagawa

Existing video-based 3D Human Mesh Recovery (HMR) methods often produce physically implausible results, stemming from their reliance on flawed intermediate 3D pose anchors and their inability to effectively model complex spatiotemporal…

Computer Vision and Pattern Recognition · Computer Science 2026-01-30 Hongjun Chen , Huan Zheng , Wencheng Han , Jianbing Shen

In this research, we address the challenge faced by existing deep learning-based human mesh reconstruction methods in balancing accuracy and computational efficiency. These methods typically prioritize accuracy, resulting in large network…

Computer Vision and Pattern Recognition · Computer Science 2023-02-01 Ayman Ali , Ekkasit Pinyoanuntapong , Pu Wang , Mohsen Dorodchi

State Space Models (SSMs) have become serious contenders in the field of sequential modeling, challenging the dominance of Transformers. At the same time, Mixture of Experts (MoE) has significantly improved Transformer-based Large Language…

This paper introduces ELMO, a real-time upsampling motion capture framework designed for a single LiDAR sensor. Modeled as a conditional autoregressive transformer-based upsampling motion generator, ELMO achieves 60 fps motion capture from…

Graphics · Computer Science 2024-12-03 Deok-Kyeong Jang , Dongseok Yang , Deok-Yun Jang , Byeoli Choi , Donghoon Shin , Sung-hee Lee

The rapid development of autonomous driving, abnormal behavior detection, and behavior recognition makes an increasing demand for multi-person pose estimation-based applications, especially on mobile platforms. However, to achieve high…

Computer Vision and Pattern Recognition · Computer Science 2021-06-30 Xuan Shen , Geng Yuan , Wei Niu , Xiaolong Ma , Jiexiong Guan , Zhengang Li , Bin Ren , Yanzhi Wang

We introduce YOLO-pose, a novel heatmap-free approach for joint detection, and 2D multi-person pose estimation in an image based on the popular YOLO object detection framework. Existing heatmap based two-stage approaches are sub-optimal as…

Computer Vision and Pattern Recognition · Computer Science 2022-04-15 Debapriya Maji , Soyeb Nagori , Manu Mathew , Deepak Poddar

Stereo disparity estimation is crucial for obtaining depth information in robot-assisted minimally invasive surgery (RAMIS). While current deep learning methods have made significant advancements, challenges remain in achieving an optimal…

Computer Vision and Pattern Recognition · Computer Science 2026-05-05 Xu Wang , Jialang Xu , Shuai Zhang , Baoru Huang , Danail Stoyanov , Evangelos B. Mazomenos

Multimodal semantic learning plays a critical role in embodied intelligence, especially when robots perceive their surroundings, understand human instructions, and make intelligent decisions. However, the field faces technical challenges…

Robotics · Computer Science 2025-09-24 Zeyi Kang , Liang He , Yanxin Zhang , Zuheng Ming , Kaixing Zhao

Transformer is popular in recent 3D human pose estimation, which utilizes long-term modeling to lift 2D keypoints into the 3D space. However, current transformer-based methods do not fully exploit the prior knowledge of the human skeleton…

Computer Vision and Pattern Recognition · Computer Science 2023-06-19 Yaqi Zhang , Yan Lu , Bin Liu , Zhiwei Zhao , Qi Chu , Nenghai Yu

Exposure Correction (EC) aims to recover proper exposure conditions for images captured under over-exposure or under-exposure scenarios. While existing deep learning models have shown promising results, few have fully embedded Retinex…

Computer Vision and Pattern Recognition · Computer Science 2024-10-30 Wei Dong , Han Zhou , Yulun Zhang , Xiaohong Liu , Jun Chen

We present EgoPoseFormer, a simple yet effective transformer-based model for stereo egocentric human pose estimation. The main challenge in egocentric pose estimation is overcoming joint invisibility, which is caused by self-occlusion or a…

Computer Vision and Pattern Recognition · Computer Science 2024-08-16 Chenhongyi Yang , Anastasia Tkach , Shreyas Hampali , Linguang Zhang , Elliot J. Crowley , Cem Keskin

Recently, polarimetric synthetic aperture radar (PolSAR) image classification has been greatly promoted by deep neural networks. However,current deep learning-based PolSAR classification methods encounter difficulties due to its dependence…

Computer Vision and Pattern Recognition · Computer Science 2025-06-03 Zuzheng Kuang , Haixia Bi , Chen Xu , Jian Sun

In recent years, 2D-to-3D pose uplifting in monocular 3D Human Pose Estimation (HPE) has attracted widespread research interest. GNN-based methods and Transformer-based methods have become mainstream architectures due to their advanced…

Computer Vision and Pattern Recognition · Computer Science 2024-10-23 Mengmeng Cui , Kunbo Zhang , Zhenan Sun

Head avatars animated by visual signals have gained popularity, particularly in cross-driving synthesis where the driver differs from the animated character, a challenging but highly practical approach. The recently presented MegaPortraits…

Computer Vision and Pattern Recognition · Computer Science 2024-05-01 Nikita Drobyshev , Antoni Bigata Casademunt , Konstantinos Vougioukas , Zoe Landgraf , Stavros Petridis , Maja Pantic

Micro-expression recognition can obtain the real emotion of the individual at the current moment. Although deep learning-based methods, especially Transformer-based methods, have achieved impressive results, these methods have high…

Computer Vision and Pattern Recognition · Computer Science 2026-04-10 Junbo Wang , Liangyu Fu , Yuke Li , Yining Zhu , Xuecheng Wu , Kun Hu

In this work, we tackle the challenge of enhancing the realism and expressiveness in talking head video generation by focusing on the dynamic and nuanced relationship between audio cues and facial movements. We identify the limitations of…

Computer Vision and Pattern Recognition · Computer Science 2024-08-09 Linrui Tian , Qi Wang , Bang Zhang , Liefeng Bo

Mamba-based models have recently demonstrated significant potential in hyperspectral image (HSI) classification, primarily due to their ability to perform contextual modeling with linear computational complexity. However, existing…

Computer Vision and Pattern Recognition · Computer Science 2025-10-07 Yichu Xu , Di Wang , Hongzan Jiao , Lefei Zhang , Liangpei Zhang