中文
相关论文

相关论文: D3RoMa: Disparity Diffusion-based Depth Sensing fo…

200 篇论文

We present a novel approach for estimating depth from a monocular camera as it moves through complex and crowded indoor environments, e.g., a department store or a metro station. Our approach predicts absolute scale depth maps over the…

计算机视觉与模式识别 · 计算机科学 2021-08-13 Dongki Jung , Jaehoon Choi , Yonghan Lee , Deokhwa Kim , Changick Kim , Dinesh Manocha , Donghwan Lee

Current discriminative depth estimation methods often produce blurry artifacts, while generative approaches suffer from slow sampling due to curvatures in the noise-to-depth transport. Our method addresses these challenges by framing depth…

Estimating precise metric depth and scene reconstruction from monocular endoscopy is a fundamental task for surgical navigation in robotic surgery. However, traditional stereo matching adopts binocular images to perceive the depth…

机器人学 · 计算机科学 2022-11-29 Ruofeng Wei , Bin Li , Hangjie Mo , Fangxun Zhong , Yonghao Long , Qi Dou , Yun-Hui Liu , Dong Sun

Depth estimation is a core task in 3D computer vision. Recent methods investigate the task of monocular depth trained with various depth sensor modalities. Every sensor has its advantages and drawbacks caused by the nature of estimates. In…

Text-to-3D generation has shown rapid progress in recent days with the advent of score distillation, a methodology of using pretrained text-to-2D diffusion models to optimize neural radiance field (NeRF) in the zero-shot setting. However,…

计算机视觉与模式识别 · 计算机科学 2024-02-07 Junyoung Seo , Wooseok Jang , Min-Seop Kwak , Hyeonsu Kim , Jaehoon Ko , Junho Kim , Jin-Hwa Kim , Jiyoung Lee , Seungryong Kim

We introduce ThermoStereoRT, a real-time thermal stereo matching method designed for all-weather conditions that recovers disparity from two rectified thermal stereo images, envisioning applications such as night-time drone surveillance or…

计算机视觉与模式识别 · 计算机科学 2025-04-11 Anning Hu , Ang Li , Xirui Jin , Danping Zou

Depth from a monocular video can enable billions of devices and robots with a single camera to see the world in 3D. In this paper, we present an approach with a differentiable flow-to-depth layer for video depth estimation. The model…

计算机视觉与模式识别 · 计算机科学 2020-03-04 Jiaxin Xie , Chenyang Lei , Zhuwen Li , Li Erran Li , Qifeng Chen

Modern cameras with large apertures often suffer from a shallow depth of field, resulting in blurry images of objects outside the focal plane. This limitation is particularly problematic for fixed-focus cameras, such as those used in smart…

计算机视觉与模式识别 · 计算机科学 2025-10-28 Xinge Yang , Chuong Nguyen , Wenbin Wang , Kaizhang Kang , Wolfgang Heidrich , Xiaoxing Li

In the field of Robot Learning, the complex mapping between high-dimensional observations such as RGB images and low-level robotic actions, two inherently very different spaces, constitutes a complex learning problem, especially with…

机器人学 · 计算机科学 2024-05-29 Vitalis Vosylius , Younggyo Seo , Jafar Uruç , Stephen James

Though the background is an important signal for image classification, over reliance on it can lead to incorrect predictions when spurious correlations between foreground and background are broken at test time. Training on a dataset where…

计算机视觉与模式识别 · 计算机科学 2022-11-21 Priyatham Kattakinda , Alexander Levine , Soheil Feizi

Learning-based methods to solve dense 3D vision problems typically train on 3D sensor data. The respectively used principle of measuring distances provides advantages and drawbacks. These are typically not compared nor discussed in the…

We propose Stereo Direct Sparse Odometry (Stereo DSO) as a novel method for highly accurate real-time visual odometry estimation of large-scale environments from stereo cameras. It jointly optimizes for all the model parameters within the…

计算机视觉与模式识别 · 计算机科学 2017-08-29 Rui Wang , Martin Schwörer , Daniel Cremers

Modeling radio frequency (RF) signal propagation is essential for understanding the environment, as RF signals offer valuable insights beyond the capabilities of RGB cameras, which are limited by the visible-light spectrum, lens coverage,…

机器学习 · 计算机科学 2025-10-07 Kyoungjun Park , Yifan Yang , Changhan Ge , Lili Qiu , Shiqi Jiang

Time-of-Flight (ToF) sensors efficiently capture scene depth, but the nonlinear depth construction procedure often results in extremely large noise variance or even invalid areas. Recent methods based on deep neural networks (DNNs) achieve…

计算机视觉与模式识别 · 计算机科学 2025-03-26 Changyong He , Jin Zeng , Jiawei Zhang , Jiajie Guo

Robotic-assisted surgery allows surgeons to conduct precise surgical operations with stereo vision and flexible motor control. However, the lack of 3D spatial perception limits situational awareness during procedures and hinders mastering…

图像与视频处理 · 电气工程与系统科学 2022-03-07 Shang Zhao , Ce Wang , Qiyuan Wang , Yanzhe Liu , S Kevin Zhou

Deformable image registration aims to precisely align medical images from different modalities or times. Traditional deep learning methods, while effective, often lack interpretability, real-time observability and adjustment capacity during…

计算机视觉与模式识别 · 计算机科学 2024-10-08 Yongtai Zhuo , Yiqing Shen

Despite the increasing prevalence of rotating-style capture (e.g., surveillance cameras), conventional stereo rectification techniques frequently fail due to the rotation-dominant motion and small baseline between views. In this paper, we…

计算机视觉与模式识别 · 计算机科学 2023-07-12 Yongcong Zhang , Yifei Xue , Ming Liao , Huiqing Zhang , Yizhen Lao

Monocular depth estimation, enabled by self-supervised learning, is a key technique for 3D perception in computer vision. However, it faces significant challenges in real-world scenarios, which encompass adverse weather variations, motion…

计算机视觉与模式识别 · 计算机科学 2024-10-10 Runze Chen , Haiyong Luo , Fang Zhao , Jingze Yu , Yupeng Jia , Juan Wang , Xuepeng Ma

Transparent objects are common in daily life. However, depth sensing for transparent objects remains a challenging problem. While learning-based methods can leverage shape priors to improve the sensing quality, the labor-intensive data…

机器人学 · 计算机科学 2023-09-19 Liuyu Bian , Pengyang Shi , Weihang Chen , Jing Xu , Li Yi , Rui Chen

We propose MonoSE(3)-Diffusion, a monocular SE(3) diffusion framework that formulates markerless, image-based robot pose estimation as a conditional denoising diffusion process. The framework consists of two processes: a…

计算机视觉与模式识别 · 计算机科学 2025-10-14 Kangjian Zhu , Haobo Jiang , Yigong Zhang , Jianjun Qian , Jian Yang , Jin Xie