English
Related papers

Related papers: SkelSplat: Robust Multi-view 3D Human Pose Estimat…

200 papers

Recognizing arbitrary or previously unseen categories is essential for comprehensive real-world 3D scene understanding. Currently, all existing methods rely on 2D or textual modalities during training or together at inference. This…

Computer Vision and Pattern Recognition · Computer Science 2025-06-04 Yue Li , Qi Ma , Runyi Yang , Huapeng Li , Mengjiao Ma , Bin Ren , Nikola Popovic , Nicu Sebe , Ender Konukoglu , Theo Gevers , Luc Van Gool , Martin R. Oswald , Danda Pani Paudel

The rapid development of multi-view 3D human pose estimation (HPE) is attributed to the maturation of monocular 2D HPE and the geometry of 3D reconstruction. However, 2D detection outliers in occluded views due to neglect of view…

Computer Vision and Pattern Recognition · Computer Science 2023-02-24 Xiaoyue Wan , Zhuo Chen , Xu Zhao

Novel view synthesis for dynamic scenes is still a challenging problem in computer vision and graphics. Recently, Gaussian splatting has emerged as a robust technique to represent static scenes and enable high-quality and real-time novel…

Computer Vision and Pattern Recognition · Computer Science 2024-04-15 Yi-Hua Huang , Yang-Tian Sun , Ziyi Yang , Xiaoyang Lyu , Yan-Pei Cao , Xiaojuan Qi

While existing feed-forward Gaussian splatting models offer computational efficiency and can generalize to sparse view settings, their performance is fundamentally constrained by relying on a single forward pass for inference. We propose…

Computer Vision and Pattern Recognition · Computer Science 2026-03-13 Haofei Xu , Daniel Barath , Andreas Geiger , Marc Pollefeys

Digitizing 3D static scenes and 4D dynamic events from multi-view images has long been a challenge in computer vision and graphics. Recently, 3D Gaussian Splatting (3DGS) has emerged as a practical and scalable reconstruction method,…

Computer Vision and Pattern Recognition · Computer Science 2024-09-18 Marko Mihajlovic , Sergey Prokudin , Siyu Tang , Robert Maier , Federica Bogo , Tony Tung , Edmond Boyer

Visual relocalization is a fundamental task in the field of 3D computer vision, estimating a camera's pose when it revisits a previously known scene. While point-based hierarchical relocalization methods have shown strong scalability and…

Computer Vision and Pattern Recognition · Computer Science 2026-04-01 Huaqi Tao , Bingxi Liu , Guangcheng Chen , Fulin Tang , Li He , Hong Zhang

We introduce SkelFormer, a novel markerless motion capture pipeline for multi-view human pose and shape estimation. Our method first uses off-the-shelf 2D keypoint estimators, pre-trained on large-scale in-the-wild data, to obtain 3D joint…

Computer Vision and Pattern Recognition · Computer Science 2024-04-22 Vandad Davoodnia , Saeed Ghorbani , Alexandre Messier , Ali Etemad

6-DoF pose estimation is a fundamental task in computer vision with wide-ranging applications in augmented reality and robotics. Existing single RGB-based methods often compromise accuracy due to their reliance on initial pose estimates and…

Computer Vision and Pattern Recognition · Computer Science 2025-03-10 Linqi Yang , Xiongwei Zhao , Qihao Sun , Ke Wang , Ao Chen , Peng Kang

Supervised approaches to 3D pose estimation from single images are remarkably effective when labeled data is abundant. However, as the acquisition of ground-truth 3D labels is labor intensive and time consuming, recent attention has shifted…

Computer Vision and Pattern Recognition · Computer Science 2022-06-30 Soumava Kumar Roy , Leonardo Citraro , Sina Honari , Pascal Fua

This paper studies the task of estimating the 3D human poses of multiple persons from multiple calibrated camera views. Following the top-down paradigm, we decompose the task into two stages, i.e. person localization and pose estimation.…

Computer Vision and Pattern Recognition · Computer Science 2021-09-14 Size Wu , Sheng Jin , Wentao Liu , Lei Bai , Chen Qian , Dong Liu , Wanli Ouyang

Talking Head Generation aims at synthesizing natural-looking talking videos from speech and a single portrait image. Previous 3D talking head generation methods have relied on domain-specific heuristics such as warping-based facial motion…

Computer Vision and Pattern Recognition · Computer Science 2026-01-27 Tong Shi , Melonie de Almeida , Daniela Ivanova , Nicolas Pugeault , Paul Henderson

3D Gaussians have recently emerged as a highly efficient representation for 3D reconstruction and rendering. Despite its high rendering quality and speed at high resolutions, they both deteriorate drastically when rendered at lower…

Computer Vision and Pattern Recognition · Computer Science 2024-05-30 Zhiwen Yan , Weng Fei Low , Yu Chen , Gim Hee Lee

3D affordance reasoning, the task of associating human instructions with the functional regions of 3D objects, is a critical capability for embodied agents. Current methods based on 3D Gaussian Splatting (3DGS) are fundamentally limited to…

Computer Vision and Pattern Recognition · Computer Science 2025-08-01 Di Li , Jie Feng , Jiahao Chen , Weisheng Dong , Guanbin Li , Yuhui Zheng , Mingtao Feng , Guangming Shi

We introduce pixelSplat, a feed-forward model that learns to reconstruct 3D radiance fields parameterized by 3D Gaussian primitives from pairs of images. Our model features real-time and memory-efficient rendering for scalable training as…

Computer Vision and Pattern Recognition · Computer Science 2024-04-08 David Charatan , Sizhe Li , Andrea Tagliasacchi , Vincent Sitzmann

Estimating 3D human poses from video is a challenging problem. The lack of 3D human pose annotations is a major obstacle for supervised training and for generalization to unseen datasets. In this work, we address this problem by proposing a…

Computer Vision and Pattern Recognition · Computer Science 2021-05-17 Mohsen Gholami , Ahmad Rezaei , Helge Rhodin , Rabab Ward , Z. Jane Wang

Despite significant advancements in dynamic neural rendering, existing methods fail to address the unique challenges posed by UAV-captured scenarios, particularly those involving monocular camera setups, top-down perspective, and multiple…

Computer Vision and Pattern Recognition · Computer Science 2025-06-06 Jaehoon Choi , Dongki Jung , Christopher Maxey , Yonghan Lee , Sungmin Eum , Dinesh Manocha , Heesung Kwon

Although monocular 3D human pose estimation methods have made significant progress, it is far from being solved due to the inherent depth ambiguity. Instead, exploiting multi-view information is a practical way to achieve absolute 3D human…

Computer Vision and Pattern Recognition · Computer Science 2022-05-18 Guoliang Hua , Hong Liu , Wenhao Li , Qian Zhang , Runwei Ding , Xin Xu

Estimating 3D human poses from 2D images remains challenging due to occlusions and projective ambiguity. Multi-view learning-based approaches mitigate these issues but often fail to generalize to real-world scenarios, as large-scale…

Computer Vision and Pattern Recognition · Computer Science 2025-12-18 Seyed Abolfazl Ghasemzadeh , Alexandre Alahi , Christophe De Vleeschouwer

Accurate 3D reconstruction in degraded imaging conditions remains a key challenge in photogrammetry and neural rendering. In underwater environments, spatially varying visibility caused by scattering, attenuation, and sparse observations…

Computer Vision and Pattern Recognition · Computer Science 2026-04-23 Zhuodong Jiang , Haoran Wang , Guoxi Huang , Brett Seymour , Nantheera Anantrasirichai

We propose NEDS-SLAM, a dense semantic SLAM system based on 3D Gaussian representation, that enables robust 3D semantic mapping, accurate camera tracking, and high-quality rendering in real-time. In the system, we propose a Spatially…

Computer Vision and Pattern Recognition · Computer Science 2024-09-04 Yiming Ji , Yang Liu , Guanghu Xie , Boyu Ma , Zongwu Xie
‹ Prev 1 8 9 10 Next ›