中文
相关论文

相关论文: PrimeDepth: Efficient Monocular Depth Estimation w…

200 篇论文

This paper presents Pixel-Perfect Depth, a monocular depth estimation model based on pixel-space diffusion generation that produces high-quality, flying-pixel-free point clouds from estimated depth maps. Current generative depth estimation…

计算机视觉与模式识别 · 计算机科学 2025-10-30 Gangwei Xu , Haotong Lin , Hongcheng Luo , Xianqi Wang , Jingfeng Yao , Lianghui Zhu , Yuechuan Pu , Cheng Chi , Haiyang Sun , Bing Wang , Guang Chen , Hangjun Ye , Sida Peng , Xin Yang

This work presents EndoStreamDepth, a monocular depth estimation framework for endoscopic video streams. It provides accurate depth maps with sharp anatomical boundaries for each frame, temporally consistent predictions across frames, and…

计算机视觉与模式识别 · 计算机科学 2026-01-05 Hao Li , Daiwei Lu , Jiacheng Wang , Robert J. Webster , Ipek Oguz

Blind face restoration methods have shown remarkable performance, particularly when trained on large-scale synthetic datasets with supervised learning. These datasets are often generated by simulating low-quality face images with a…

计算机视觉与模式识别 · 计算机科学 2024-11-26 Tianshu Kuai , Sina Honari , Igor Gilitschenski , Alex Levinshtein

In this paper, we propose a dense monocular SLAM system, named DeepRelativeFusion, that is capable to recover a globally consistent 3D structure. To this end, we use a visual SLAM algorithm to reliably recover the camera poses and…

计算机视觉与模式识别 · 计算机科学 2021-07-13 Shing Yan Loo , Syamsiah Mashohor , Sai Hong Tang , Hong Zhang

Monocular depth estimation is an ambiguous problem, thus global structural cues play an important role in current data-driven single-view depth estimation methods. Panorama images capture the complete spatial information of their…

计算机视觉与模式识别 · 计算机科学 2023-01-18 Meng Li , Senbo Wang , Weihao Yuan , Weichao Shen , Zhe Sheng , Zilong Dong

The problem of depth completion involves predicting a dense depth image from a single sparse depth map and an RGB image. Unsupervised depth completion methods have been proposed for various datasets where ground truth depth data is…

计算机视觉与模式识别 · 计算机科学 2025-12-16 Sangmin Hong , Suyoung Lee , Kyoung Mu Lee

Extensive pre-training with large data is indispensable for downstream geometry and semantic visual perception tasks. Thanks to large-scale text-to-image (T2I) pretraining, recent works show promising results by simply fine-tuning T2I…

计算机视觉与模式识别 · 计算机科学 2024-12-03 Guangkai Xu , Yongtao Ge , Mingyu Liu , Chengxiang Fan , Kangyang Xie , Zhiyue Zhao , Hao Chen , Chunhua Shen

Object pose estimation is a core means for robots to understand and interact with their environment. For this task, monocular category-level methods are attractive as they require only a single RGB camera. However, current methods rely on…

计算机视觉与模式识别 · 计算机科学 2025-04-15 Jian Liu , Wei Sun , Hui Yang , Jin Zheng , Zichen Geng , Hossein Rahmani , Ajmal Mian

We present Buffer Anytime, a framework for estimation of depth and normal maps (which we call geometric buffers) from video that eliminates the need for paired video--depth and video--normal training data. Instead of relying on large-scale…

计算机视觉与模式识别 · 计算机科学 2024-11-27 Zhengfei Kuang , Tianyuan Zhang , Kai Zhang , Hao Tan , Sai Bi , Yiwei Hu , Zexiang Xu , Milos Hasan , Gordon Wetzstein , Fujun Luan

Perceiving 3D information is of paramount importance in many applications of computer vision. Recent advances in monocular depth estimation have shown that gaining such knowledge from a single camera input is possible by training deep…

计算机视觉与模式识别 · 计算机科学 2021-10-28 Sai Shyam Chanduri , Zeeshan Khan Suri , Igor Vozniak , Christian Müller

Stereo foundation models achieve strong zero-shot generalization but remain computationally prohibitive for real-time applications. Efficient stereo architectures, on the other hand, sacrifice robustness for speed and require costly…

计算机视觉与模式识别 · 计算机科学 2026-03-18 Bowen Wen , Shaurya Dewan , Stan Birchfield

We explore the problem of computationally generating special `prime' images that produce optical illusions when physically arranged and viewed in a certain way. First, we propose a formal definition for this problem. Next, we introduce…

计算机视觉与模式识别 · 计算机科学 2023-12-08 Ryan Burgert , Xiang Li , Abe Leite , Kanchana Ranasinghe , Michael S. Ryoo

Single image depth estimation is a challenging problem. The current state-of-the-art method formulates the problem as that of ordinal regression. However, the formulation is not fully differentiable and depth maps are not generated in an…

计算机视觉与模式识别 · 计算机科学 2020-06-17 Kunal Swami , Prasanna Vishnu Bondada , Pankaj Kumar Bajpai

In this paper, we propose \textbf{Iris}, a deterministic framework for Monocular Depth Estimation (MDE) that integrates real-world priors into the diffusion model. Conventional feed-forward methods rely on massive training data, yet still…

计算机视觉与模式识别 · 计算机科学 2026-03-20 Xinhao Cai , Gensheng Pei , Zeren Sun , Yazhou Yao , Fumin Shen , Wenguan Wang

The panorama image can simultaneously demonstrate complete information of the surrounding environment and has many advantages in virtual tourism, games, robotics, etc. However, the progress of panorama depth estimation cannot completely…

计算机视觉与模式识别 · 计算机科学 2022-12-06 Qingsong Yan , Qiang Wang , Kaiyong Zhao , Bo Li , Xiaowen Chu , Fei Deng

Diffusion models are a new class of generative models, and have dramatically promoted image generation with unprecedented quality and diversity. Existing diffusion models mainly try to reconstruct input image from a corrupted one with a…

计算机视觉与模式识别 · 计算机科学 2024-06-05 Ling Yang , Jingwei Liu , Shenda Hong , Zhilong Zhang , Zhilin Huang , Zheming Cai , Wentao Zhang , Bin Cui

Denoising diffusion probabilistic models (DDPMs) have recently achieved leading performances in many generative tasks. However, the inherited iterative sampling process costs hinder their applications to text-to-speech deployment. Through…

音频与语音处理 · 电气工程与系统科学 2022-07-14 Rongjie Huang , Zhou Zhao , Huadai Liu , Jinglin Liu , Chenye Cui , Yi Ren

Accurate depth estimation is at the core of many applications in computer graphics, vision, and robotics. Current state-of-the-art monocular depth estimators, trained on extensive datasets, generalize well but lack 3D consistency needed for…

计算机视觉与模式识别 · 计算机科学 2025-11-26 Laura Fink , Linus Franke , Bernhard Egger , Joachim Keinert , Marc Stamminger

Monocular depth estimation is a challenging task in complex compositions depicting multiple objects of diverse scales. Albeit the recent great progress thanks to the deep convolutional neural networks (CNNs), the state-of-the-art monocular…

计算机视觉与模式识别 · 计算机科学 2017-08-09 Bo Li , Yuchao Dai , Mingyi He

Metric depth estimation from visual sensors is crucial for robots to perceive, navigate, and interact with their environment. Traditional range imaging setups, such as stereo or structured light cameras, face hassles including calibration,…

计算机视觉与模式识别 · 计算机科学 2025-03-13 Blanca Lasheras-Hernandez , Klaus H. Strobl , Sergio Izquierdo , Tim Bodenmüller , Rudolph Triebel , Javier Civera