中文
相关论文

相关论文: HAMSt3R: Human-Aware Multi-view Stereo 3D Reconstr…

200 篇论文

This paper addresses the task of large-scale 3D scene reconstruction from long video sequences. Recent feed-forward reconstruction models have shown promising results by directly regressing 3D geometry from RGB images without explicit 3D…

计算机视觉与模式识别 · 计算机科学 2026-04-10 Tao Xie , Peishan Yang , Yudong Jin , Yingfeng Cai , Wei Yin , Weiqiang Ren , Qian Zhang , Wei Hua , Sida Peng , Xiaoyang Guo , Xiaowei Zhou

In this paper, we present a novel framework designed to reconstruct long-sequence 3D human motion in the world coordinates from in-the-wild videos with multiple shot transitions. Such long-sequence in-the-wild motions are highly valuable to…

计算机视觉与模式识别 · 计算机科学 2025-03-11 Yuhong Zhang , Guanlin Wu , Ling-Hao Chen , Zhuokai Zhao , Jing Lin , Xiaoke Jiang , Jiamin Wu , Zhuoheng Li , Hao Frank Yang , Haoqian Wang , Lei Zhang

Human Mesh Recovery (HMR) aims to reconstruct 3D human pose and shape from 2D observations and is fundamental to human-centric understanding in real-world scenarios. While recent image-based HMR methods such as SAM 3D Body achieve strong…

计算机视觉与模式识别 · 计算机科学 2025-12-10 Mingqi Gao , Yunqi Miao , Jungong Han

Dense 3D reconstruction has many applications in automated driving including automated annotation validation, multimodal data augmentation, providing ground truth annotations for systems lacking LiDAR, as well as enhancing auto-labeling…

计算机视觉与模式识别 · 计算机科学 2024-02-13 Shihao Shen , Louis Kerofsky , Varun Ravi Kumar , Senthil Yogamani

Streaming 3D perception is well suited to robotics and augmented reality, where long visual streams must be processed efficiently and consistently. Recent recurrent models offer a promising solution by maintaining fixed-size states and…

计算机视觉与模式识别 · 计算机科学 2026-04-09 Changkun Liu , Jiezhi Yang , Zeman Li , Yuan Deng , Jiancong Guo , Luca Ballan

Humans can naturally identify and mentally complete occluded objects in cluttered environments. However, imparting similar cognitive ability to robotics remains challenging even with advanced reconstruction techniques, which models scenes…

计算机视觉与模式识别 · 计算机科学 2025-07-22 Zesong Yang , Bangbang Yang , Wenqi Dong , Chenxuan Cao , Liyuan Cui , Yuewen Ma , Zhaopeng Cui , Hujun Bao

Robots often rely on RGB images for tasks like manipulation and navigation. However, reliable interaction typically requires a 3D scene representation that is metric-scaled and aligned with the robot reference frame. This depends on…

机器人学 · 计算机科学 2025-09-11 Davide Allegro , Matteo Terreran , Stefano Ghidoni

Head-mounted displays (HMDs) are essential for experiencing extended reality (XR) environments and observing virtual content. However, they obscure the upper part of the user's face, complicating external video recording and significantly…

计算机视觉与模式识别 · 计算机科学 2025-08-19 Fatemeh Ghorbani Lohesara , Karen Eguiazarian , Sebastian Knorr

We introduce HART, a unified framework for sparse-view human reconstruction. Given a small set of uncalibrated RGB images of a person as input, it outputs a watertight clothed mesh, the aligned SMPL-X body mesh, and a Gaussian-splat…

计算机视觉与模式识别 · 计算机科学 2025-10-01 Xiyi Chen , Shaofei Wang , Marko Mihajlovic , Taewon Kang , Sergey Prokudin , Ming Lin

Reconstructing 3D human shape and pose from monocular images is challenging despite the promising results achieved by the most recent learning-based methods. The commonly occurred misalignment comes from the facts that the mapping from…

计算机视觉与模式识别 · 计算机科学 2020-12-08 Hongwen Zhang , Jie Cao , Guo Lu , Wanli Ouyang , Zhenan Sun

We present THUNDR, a transformer-based deep neural network methodology to reconstruct the 3d pose and shape of people, given monocular RGB images. Key to our methodology is an intermediate 3d marker representation, where we aim to combine…

计算机视觉与模式识别 · 计算机科学 2021-06-18 Mihai Zanfir , Andrei Zanfir , Eduard Gabriel Bazavan , William T. Freeman , Rahul Sukthankar , Cristian Sminchisescu

While recent feed-forward 3D reconstruction models accelerate 3D reconstruction by jointly inferring dense geometry and camera poses in a single pass, their reliance on dense attention imposes a quadratic complexity, creating a prohibitive…

计算机视觉与模式识别 · 计算机科学 2026-03-10 Weining Ren , Xiao Tan , Kai Han

Transformer-based 3D reconstruction has emerged as a powerful paradigm for recovering geometry and appearance from multi-view observations, offering strong performance across challenging visual conditions. As these models scale to larger…

计算机视觉与模式识别 · 计算机科学 2026-05-13 Haoyu Zhang , Zeyu Zhang , Zedong Zhou , Yang Zhao , Hao Tang

Reconstructing detailed 3D scenes from single-view images remains a challenging task due to limitations in existing approaches, which primarily focus on geometric shape recovery, overlooking object appearances and fine shape details. To…

计算机视觉与模式识别 · 计算机科学 2023-11-02 Yixin Chen , Junfeng Ni , Nan Jiang , Yaowei Zhang , Yixin Zhu , Siyuan Huang

Performing single image holistic understanding and 3D reconstruction is a central task in computer vision. This paper presents an integrated system that performs dense scene labeling, object detection, instance segmentation, depth…

计算机视觉与模式识别 · 计算机科学 2021-12-01 Sainan Liu , Vincent Nguyen , Yuan Gao , Subarna Tripathi , Zhuowen Tu

3D object reconstruction and multilevel segmentation are fundamental to computer vision research. Existing algorithms usually perform 3D scene reconstruction and target objects segmentation independently, and the performance is not fully…

计算机视觉与模式识别 · 计算机科学 2023-08-15 Jiexiong Xu , Weikun Zhao , Zhiyan Tang , Xiangchao Gan

Full 3D estimation of human pose from a single image remains a challenging task despite many recent advances. In this paper, we explore the hypothesis that strong prior information about scene geometry can be used to improve pose estimation…

计算机视觉与模式识别 · 计算机科学 2021-12-10 Zhe Wang , Liyan Chen , Shaurya Rathore , Daeyun Shin , Charless Fowlkes

This paper presents a novel framework to recover detailed human body shapes from a single image. It is a challenging task due to factors such as variations in human shapes, body poses, and viewpoints. Prior methods typically attempt to…

计算机视觉与模式识别 · 计算机科学 2019-05-14 Hao Zhu , Xinxin Zuo , Sen Wang , Xun Cao , Ruigang Yang

We propose a computational framework to jointly parse a single RGB image and reconstruct a holistic 3D configuration composed by a set of CAD models using a stochastic grammar model. Specifically, we introduce a Holistic Scene Grammar (HSG)…

计算机视觉与模式识别 · 计算机科学 2018-08-08 Siyuan Huang , Siyuan Qi , Yixin Zhu , Yinxue Xiao , Yuanlu Xu , Song-Chun Zhu

Recent progress in human shape learning, shows that neural implicit models are effective in generating 3D human surfaces from limited number of views, and even from a single RGB image. However, existing monocular approaches still struggle…

计算机视觉与模式识别 · 计算机科学 2024-03-20 Marco Pesavento , Yuanlu Xu , Nikolaos Sarafianos , Robert Maier , Ziyan Wang , Chun-Han Yao , Marco Volino , Edmond Boyer , Adrian Hilton , Tony Tung