English
Related papers

Related papers: GemDepth: Geometry-Embedded Features for 3D-Consis…

200 papers

Complex geometric tasks such as geometric modeling, physical simulation, and texture parametrization often involve the embedding of many complex sub-domains with potentially different dimensions. These tasks often require evolving the…

Graphics · Computer Science 2025-01-03 Michael Tao , Jiacheng Dai , Denis Zorin , Teseo Schneider , Daniele Panozzo

Reconstructing dynamic 4D scenes is an important yet challenging task. While 3D foundation models like VGGT excel in static settings, they often struggle with dynamic sequences where motion causes significant geometric ambiguity. To address…

Computer Vision and Pattern Recognition · Computer Science 2026-04-13 Ying Zang , Yidong Han , Chaotao Ding , Yuanqi Hu , Deyi Ji , Qi Zhu , Xuanfu Li , Jin Ma , Lingyun Sun , Tianrun Chen , Lanyun Zhu

Depth estimation is a cornerstone of 3D reconstruction and plays a vital role in minimally invasive endoscopic surgeries. However, most current depth estimation networks rely on traditional convolutional neural networks, which are limited…

Computer Vision and Pattern Recognition · Computer Science 2025-07-16 Bojian Li , Bo Liu , Xinning Yao , Jinghua Yue , Fugen Zhou

We present a novel method for multi-view depth estimation from a single video, which is a critical task in various applications, such as perception, reconstruction and robot navigation. Although previous learning-based methods have…

Computer Vision and Pattern Recognition · Computer Science 2021-07-13 Xiaoxiao Long , Lingjie Liu , Wei Li , Christian Theobalt , Wenping Wang

The 3D video quality metric (3VQM) was proposed to evaluate the temporal and spatial variation of the depth errors for the depth values that would lead to inconsistencies between left and right views, fast changing disparities, and…

Image and Video Processing · Electrical Eng. & Systems 2018-11-22 Dogancan Temel , Ghassan AlRegib

Geometric priors are often used to enhance 3D reconstruction. With many smartphones featuring low-resolution depth sensors and the prevalence of off-the-shelf monocular geometry estimators, incorporating geometric priors as regularization…

Computer Vision and Pattern Recognition · Computer Science 2024-12-17 Xuqian Ren , Matias Turkulainen , Jiepeng Wang , Otto Seiskari , Iaroslav Melekhov , Juho Kannala , Esa Rahtu

We propose a depth map inference system from monocular videos based on a novel dataset for navigation that mimics aerial footage from gimbal stabilized monocular camera in rigid scenes. Unlike most navigation datasets, the lack of rotation…

Computer Vision and Pattern Recognition · Computer Science 2018-09-13 Clément Pinard , Laure Chevalley , Antoine Manzanera , David Filliat

We propose an online multi-view depth prediction approach on posed video streams, where the scene geometry information computed in the previous time steps is propagated to the current time step in an efficient and geometrically plausible…

Computer Vision and Pattern Recognition · Computer Science 2021-07-23 Arda Düzçeker , Silvano Galliani , Christoph Vogel , Pablo Speciale , Mihai Dusmanu , Marc Pollefeys

Significant advancements have been made in video generative models recently. Unlike image generation, video generation presents greater challenges, requiring not only generating high-quality frames but also ensuring temporal consistency…

Computer Vision and Pattern Recognition · Computer Science 2024-07-24 Jiahe Liu , Youran Qu , Qi Yan , Xiaohui Zeng , Lele Wang , Renjie Liao

Monocular depth estimation is scale-ambiguous, and thus requires scale supervision to produce metric predictions. Even so, the resulting models will be geometry-specific, with learned scales that cannot be directly transferred across…

Computer Vision and Pattern Recognition · Computer Science 2023-07-03 Vitor Guizilini , Igor Vasiljevic , Dian Chen , Rares Ambrus , Adrien Gaidon

Understanding the 3D geometry and semantics of driving scenes is critical for safe autonomous driving. Recent advances in 3D occupancy prediction have improved scene representation but often suffer from visual inconsistencies, leading to…

Computer Vision and Pattern Recognition · Computer Science 2025-07-08 Loïck Chambon , Eloi Zablocki , Alexandre Boulch , Mickaël Chen , Matthieu Cord

With the advancement of computer vision, dynamic 3D reconstruction techniques have seen significant progress and found applications in various fields. However, these techniques generate large amounts of 3D data sequences, necessitating…

Computer Vision and Pattern Recognition · Computer Science 2024-09-10 Haichao Zhu

Current geometry-based monocular 3D object detection models can efficiently detect objects by leveraging perspective geometry, but their performance is limited due to the absence of accurate depth information. Though this issue can be…

Computer Vision and Pattern Recognition · Computer Science 2021-07-29 Chenhang He , Jianqiang Huang , Xian-Sheng Hua , Lei Zhang

Gaussian Splatting has been considered as a novel way for view synthesis of dynamic scenes, which shows great potential in AIoT applications such as digital twins. However, recent dynamic Gaussian Splatting methods significantly degrade…

Computer Vision and Pattern Recognition · Computer Science 2025-12-01 Yiwei Li , Jiannong Cao , Penghui Ruan , Divya Saxena , Songye Zhu , Yinfeng Cao

This paper performs the first investigation into depth for large-scale human action recognition in video where the depth cues are estimated from the videos themselves. We develop a new framework called depth2action and experiment thoroughly…

Computer Vision and Pattern Recognition · Computer Science 2016-08-16 Yi Zhu , Shawn Newsam

Accurate and temporally consistent modeling of human bodies is essential for a wide range of applications, including character animation, understanding human social behavior and AR/VR interfaces. Capturing human motion accurately from a…

Computer Vision and Pattern Recognition · Computer Science 2022-02-09 Alexandra Zimmer , Anna Hilsmann , Wieland Morgenstern , Peter Eisert

Geometry plays a significant role in monocular 3D object detection. It can be used to estimate object depth by using the perspective projection between object's physical size and 2D projection in the image plane, which can introduce…

Computer Vision and Pattern Recognition · Computer Science 2025-01-08 Yan Lu , Xinzhu Ma , Lei Yang , Tianzhu Zhang , Yating Liu , Qi Chu , Tong He , Yonghui Li , Wanli Ouyang

In this paper, we present a fast monocular depth estimation method for enabling 3D perception capabilities of low-cost underwater robots. We formulate a novel end-to-end deep visual learning pipeline named UDepth, which incorporates domain…

Computer Vision and Pattern Recognition · Computer Science 2023-02-03 Boxiao Yu , Jiayi Wu , Md Jahidul Islam

Achieving human-like reasoning in deep learning models for complex tasks in unknown environments remains a critical challenge in embodied intelligence. While advanced vision-language models (VLMs) excel in static scene understanding, their…

Computer Vision and Pattern Recognition · Computer Science 2025-09-03 Jinzhou Tang , Jusheng zhang , Sidi Liu , Waikit Xiu , Qinhan Lv , Xiying Li

Recent video depth estimation methods achieve great performance by following the paradigm of image depth estimation, i.e., typically fine-tuning pre-trained video diffusion models with massive data. However, we argue that video depth…

Computer Vision and Pattern Recognition · Computer Science 2025-11-11 Haodong Li , Chen Wang , Jiahui Lei , Kostas Daniilidis , Lingjie Liu