English
Related papers

Related papers: UVRM: A Scalable 3D Reconstruction Model from Unpo…

200 papers

Reconstructing dynamic assets from video data is central to many in computer vision and graphics tasks. Existing 4D reconstruction approaches are limited by category-specific models or slow optimization-based methods. Inspired by the recent…

Computer Vision and Pattern Recognition · Computer Science 2025-03-31 Remy Sabathier , Niloy J. Mitra , David Novotny

Previous research has investigated the application of Multimodal Large Language Models (MLLMs) in understanding 3D scenes by interpreting them as videos. These approaches generally depend on comprehensive 3D data inputs, such as point…

Computer Vision and Pattern Recognition · Computer Science 2025-10-23 Duo Zheng , Shijia Huang , Yanyang Li , Liwei Wang

Single visual object tracking from an unmanned aerial vehicle (UAV) poses fundamental challenges such as object occlusion, small-scale objects, background clutter, and abrupt camera motion. To tackle these difficulties, we propose to…

Computer Vision and Pattern Recognition · Computer Science 2022-02-28 Stéphane Vujasinović , Stefan Becker , Timo Breuer , Sebastian Bullinger , Norbert Scherer-Negenborn , Michael Arens

Learning 3D human motion from 2D inputs is a fundamental task in the realms of computer vision and computer graphics. Many previous methods grapple with this inherently ambiguous task by introducing motion priors into the learning process.…

Computer Vision and Pattern Recognition · Computer Science 2024-04-16 Shuaiying Hou , Hongyu Tao , Junheng Fang , Changqing Zou , Hujun Bao , Weiwei Xu

Recently, deep learning based 3D face reconstruction methods have shown promising results in both quality and efficiency.However, training deep neural networks typically requires a large volume of data, whereas face images with ground-truth…

Computer Vision and Pattern Recognition · Computer Science 2020-04-10 Yu Deng , Jiaolong Yang , Sicheng Xu , Dong Chen , Yunde Jia , Xin Tong

Existing video self-supervised learning methods mainly rely on trimmed videos for model training. However, trimmed datasets are manually annotated from untrimmed videos. In this sense, these methods are not really self-supervised. In this…

Computer Vision and Pattern Recognition · Computer Science 2020-08-07 Dezhao Luo , Bo Fang , Yu Zhou , Yucan Zhou , Dayan Wu , Weiping Wang

This paper presents a method to reconstruct high-quality textured 3D models from single images. Current methods rely on datasets with expensive annotations; multi-view images and their camera parameters. Our method relies on GAN generated…

Computer Vision and Pattern Recognition · Computer Science 2023-05-19 Aysegul Dundar , Jun Gao , Andrew Tao , Bryan Catanzaro

Superior human pose and shape reconstruction from monocular images depends on removing the ambiguities caused by occlusions and shape variance. Recent works succeed in regression-based methods which estimate parametric models directly…

Computer Vision and Pattern Recognition · Computer Science 2021-02-01 Min Wang , Feng Qiu , Wentao Liu , Chen Qian , Xiaowei Zhou , Lizhuang Ma

3D reconstruction, which aims to recover the dense three-dimensional structure of a scene, is a cornerstone technology for numerous applications, including augmented/virtual reality, autonomous driving, and robotics. While traditional…

Computer Vision and Pattern Recognition · Computer Science 2025-07-14 Wei Zhang , Yihang Wu , Songhua Li , Wenjie Ma , Xin Ma , Qiang Li , Qi Wang

High-resolution (HR) videos play a crucial role in many computer vision applications. Although existing video restoration (VR) methods can significantly enhance video quality by exploiting temporal information across video frames, they are…

Computer Vision and Pattern Recognition · Computer Science 2025-06-09 Mary Aiyetigbo , Wanqi Yuan , Feng Luo , Nianyi Li

Neural volume rendering enables photo-realistic renderings of a human performer in free-view, a critical task in immersive VR/AR applications. But the practice is severely limited by high computational costs in the rendering process. To…

Computer Vision and Pattern Recognition · Computer Science 2023-03-06 Yue Chen , Xuan Wang , Xingyu Chen , Qi Zhang , Xiaoyu Li , Yu Guo , Jue Wang , Fei Wang

Learning model-free object pose estimation for unseen instances remains a fundamental challenge in 3D vision. Existing methods typically fall into two disjoint paradigms: category-level approaches predict absolute poses in a canonical space…

Computer Vision and Pattern Recognition · Computer Science 2026-03-25 Weihang Li , Lorenzo Garattoni , Fabien Despinoy , Nassir Navab , Benjamin Busam

This paper targets on learning-based novel view synthesis from a single or limited 2D images without the pose supervision. In the viewer-centered coordinates, we construct an end-to-end trainable conditional variational framework to…

Computer Vision and Pattern Recognition · Computer Science 2021-06-08 Xiaofeng Liu , Tong Che , Yiqun Lu , Chao Yang , Site Li , Jane You

In this work, we demonstrate that 3D poses in video can be effectively estimated with a fully convolutional model based on dilated temporal convolutions over 2D keypoints. We also introduce back-projection, a simple and effective…

Computer Vision and Pattern Recognition · Computer Science 2019-04-01 Dario Pavllo , Christoph Feichtenhofer , David Grangier , Michael Auli

Recently, building on the foundation of neural radiance field, various techniques have emerged to learn unsigned distance fields (UDF) to reconstruct 3D non-watertight models from multi-view images. Yet, a central challenge in UDF-based…

Computer Vision and Pattern Recognition · Computer Science 2024-04-17 Junkai Deng , Fei Hou , Xuhui Chen , Wencheng Wang , Ying He

While monocular 3D pose estimation seems to have achieved very accurate results on the public datasets, their generalization ability is largely overlooked. In this work, we perform a systematic evaluation of the existing methods and find…

Computer Vision and Pattern Recognition · Computer Science 2022-07-21 Jiajun Su , Chunyu Wang , Xiaoxuan Ma , Wenjun Zeng , Yizhou Wang

Vision-Language models (VLMs) have excelled in the image-domain -- especially in zero-shot settings -- thanks to the availability of vast pretraining data (i.e., paired image-text samples). However for videos, such paired data is not as…

Computer Vision and Pattern Recognition · Computer Science 2024-04-01 Kumara Kahatapitiya , Anurag Arnab , Arsha Nagrani , Michael S. Ryoo

Current large multimodal models (LMMs) face challenges in grounding, which requires the model to relate language components to visual entities. Contrary to the common practice that fine-tunes LMMs with additional grounding supervision, we…

Computer Vision and Pattern Recognition · Computer Science 2025-10-17 Shengcao Cao , Liang-Yan Gui , Yu-Xiong Wang

Depth estimation is a foundational component for 3D reconstruction in minimally invasive endoscopic surgeries. However, existing monocular depth estimation techniques often exhibit limited performance to the varying illumination and complex…

Computer Vision and Pattern Recognition · Computer Science 2025-08-26 Xinning Yao , Bo Liu , Bojian Li , Jingjing Wang , Jinghua Yue , Fugen Zhou

We propose RoHM, an approach for robust 3D human motion reconstruction from monocular RGB(-D) videos in the presence of noise and occlusions. Most previous approaches either train neural networks to directly regress motion in 3D or learn…

Computer Vision and Pattern Recognition · Computer Science 2024-04-16 Siwei Zhang , Bharat Lal Bhatnagar , Yuanlu Xu , Alexander Winkler , Petr Kadlecek , Siyu Tang , Federica Bogo