中文
相关论文

相关论文: Mono3R: Exploiting Monocular Cues for Geometric 3D…

200 篇论文

3D object detection is a fundamental and challenging task for 3D scene understanding, and the monocular-based methods can serve as an economical alternative to the stereo-based or LiDAR-based methods. However, accurately detecting objects…

计算机视觉与模式识别 · 计算机科学 2022-01-27 Zhiyu Chong , Xinzhu Ma , Hong Zhang , Yuxin Yue , Haojie Li , Zhihui Wang , Wanli Ouyang

We propose a simple yet effective approach to enhance the performance of feed-forward 3D reconstruction models. Existing methods often struggle near depth discontinuities, where standard regression losses encourage spatial averaging and…

计算机视觉与模式识别 · 计算机科学 2026-01-27 Zichen Wang , Ang Cao , Liam J. Wang , Jeong Joon Park

In this work, we propose a novel single-shot and keypoints-based framework for monocular 3D objects detection using only RGB images, called KM3D-Net. We design a fully convolutional model to predict object keypoints, dimension, and…

计算机视觉与模式识别 · 计算机科学 2020-09-03 Peixuan Li

Accurate Digital Surface Model (DSM) reconstruction from satellite imagery is critical for applications such as disaster response, urban planning, and large-scale geographic mapping. Existing approaches face a fundamental trade-off:…

计算机视觉与模式识别 · 计算机科学 2026-05-11 Qiaoyi Yang , Chaoyi Zhou , Xi Liu , Run Wang , Minghui Xu , Mert D. Pesé , Feng Luo , Yuhao Xu , Zhi-Qi Cheng , Qiushi Chen , Hairong Qi , Siyu Huang

Multi-object tracking (MOT) in monocular videos is fundamentally challenged by occlusions and depth ambiguity, issues that conventional tracking-by-detection (TBD) methods struggle to resolve owing to a lack of geometric awareness. To…

计算机视觉与模式识别 · 计算机科学 2025-08-12 Xudong Han , Pengcheng Fang , Yueying Tian , Jianhui Yu , Xiaohao Cai , Daniel Roggen , Philip Birch

We introduce UniCon3R, a unified feed-forward framework for online human-scene 4D reconstruction from monocular video. Current feed-forward human-scene reconstruction methods suffer from artifacts, where bodies float above the ground or…

计算机视觉与模式识别 · 计算机科学 2026-05-12 Tanuj Sur , Shashank Tripathi , Nikos Athanasiou , Ha Linh Nguyen , Kai Xu , Michael J. Black , Angela Yao

Monocular 3D human pose estimation remains a challenging and ill-posed problem, particularly in real-time settings and unconstrained environments. While direct imageto-3D approaches require large annotated datasets and heavy models,…

计算机视觉与模式识别 · 计算机科学 2025-07-24 Mohamed Adjel

The rapid advancement of Large Multimodal Models (LMMs) for 2D images and videos has motivated extending these models to understand 3D scenes, aiming for human-like visual-spatial intelligence. Nevertheless, achieving deep spatial…

Depth estimation plays a pivotal role in advancing human-robot interactions, especially in indoor environments where accurate 3D scene reconstruction is essential for tasks like navigation and object handling. Monocular depth estimation,…

计算机视觉与模式识别 · 计算机科学 2025-02-18 Siddiqui Muhammad Yasir , Hyunsik Ahn

Recent advances in 4D scene reconstruction have significantly improved dynamic modeling across various domains. However, existing approaches remain limited under aerial conditions with single-view capture, wide spatial range, and dynamic…

计算机视觉与模式识别 · 计算机科学 2026-02-27 Hanyang Liu , Rongjun Qin

Panoptic segmentation of 3D scenes, involving the segmentation and classification of object instances in a dense 3D reconstruction of a scene, is a challenging problem, especially when relying solely on unposed 2D images. Existing…

计算机视觉与模式识别 · 计算机科学 2025-06-27 Lojze Zust , Yohann Cabon , Juliette Marrie , Leonid Antsfeld , Boris Chidlovskii , Jerome Revaud , Gabriela Csurka

Solving the challenging problem of 3D object reconstruction from a single image appropriately gives existing technologies the ability to perform with a single monocular camera rather than requiring depth sensors. In recent years, thanks to…

计算机视觉与模式识别 · 计算机科学 2021-08-18 Guiju Ping , Mahdi Abolfazli Esfahani , Han Wang

Joint camera pose and dense geometry estimation from a set of images or a monocular video remains a challenging problem due to its computational complexity and inherent visual ambiguities. Most dense incremental reconstruction systems…

计算机视觉与模式识别 · 计算机科学 2024-04-18 Kirill Mazur , Gwangbin Bae , Andrew J. Davison

Recovering 3D full-body human pose is a challenging problem with many applications. It has been successfully addressed by motion capture systems with body worn markers and multiple cameras. In this paper, we address the more challenging…

计算机视觉与模式识别 · 计算机科学 2018-03-12 Xiaowei Zhou , Menglong Zhu , Georgios Pavlakos , Spyridon Leonardos , Kostantinos G. Derpanis , Kostas Daniilidis

Monocular Depth Estimation (MDE) enables spatial understanding, 3D reconstruction, and autonomous navigation, yet deep learning approaches often predict only relative depth without a consistent metric scale. This limitation reduces…

计算机视觉与模式识别 · 计算机科学 2025-08-27 Jiuling Zhang

Generalizing metric monocular depth estimation presents a significant challenge due to its ill-posed nature, while the entanglement between camera parameters and depth amplifies issues further, hindering multi-dataset training and zero-shot…

计算机视觉与模式识别 · 计算机科学 2025-09-26 Karlo Koledić , Luka Petrović , Ivan Marković , Ivan Petrović

We present MOSAIC-GS, a novel, fully explicit, and computationally efficient approach for high-fidelity dynamic scene reconstruction from monocular videos using Gaussian Splatting. Monocular reconstruction is inherently ill-posed due to the…

计算机视觉与模式识别 · 计算机科学 2026-01-12 Svitlana Morkva , Maximum Wilder-Smith , Michael Oechsle , Alessio Tonioni , Marco Hutter , Vaishakh Patil

This paper reports on a novel template-free monocular non-rigid surface reconstruction approach. Existing techniques using motion and deformation cues rely on multiple prior assumptions, are often computationally expensive and do not…

计算机视觉与模式识别 · 计算机科学 2017-10-18 Mohammad Dawud Ansari , Vladislav Golyanik , Didier Stricker

There have been attempts to detect 3D objects by fusion of stereo camera images and LiDAR sensor data or using LiDAR for pre-training and only monocular images for testing, but there have been less attempts to use only monocular image…

计算机视觉与模式识别 · 计算机科学 2022-09-21 Curie Kim , Ue-Hwan Kim , Jong-Hwan Kim

Recent sparse multi-view scene reconstruction advances like DUSt3R and MASt3R no longer require camera calibration and camera pose estimation. However, they only process a pair of views at a time to infer pixel-aligned pointmaps. When…

计算机视觉与模式识别 · 计算机科学 2024-12-11 Zhenggang Tang , Yuchen Fan , Dilin Wang , Hongyu Xu , Rakesh Ranjan , Alexander Schwing , Zhicheng Yan