中文
相关论文

相关论文: D4D: An RGBD diffusion model to boost monocular de…

200 篇论文

Estimating depth from RGB images is a long-standing ill-posed problem, which has been explored for decades by the computer vision, graphics, and machine learning communities. In this article, we provide a comprehensive survey of the recent…

计算机视觉与模式识别 · 计算机科学 2019-06-17 Hamid Laga

Previous RGB-D salient object detection (SOD) methods have widely adopted deep learning tools to automatically strike a trade-off between RGB and D (depth), whose key rationale is to take full advantage of their complementary nature, aiming…

计算机视觉与模式识别 · 计算机科学 2020-08-11 Xuehao Wang , Shuai Li , Chenglizhao Chen , Aimin Hao , Hong Qin

Using the raw data from consumer-level RGB-D cameras as input, we propose a deep-learning based approach to efficiently generate RGB-D images with completed information in high resolution. To process the input images in low resolution with…

计算机视觉与模式识别 · 计算机科学 2020-06-15 Chuhua Xian , Dongjiu Zhang , Chengkai Dai , Charlie C. L. Wang

In the RGB-D vision community, extensive research has been focused on designing multi-modal learning strategies and fusion structures. However, the complementary and fusion mechanisms in RGB-D models remain a black box. In this paper, we…

计算机视觉与模式识别 · 计算机科学 2025-06-17 Hao Chen , Haoran Zhou , Yunshu Zhang , Zheng Lin , Yongjian Deng

We study data-free knowledge distillation (KD) for monocular depth estimation (MDE), which learns a lightweight model for real-world depth perception tasks by compressing it from a trained teacher model while lacking training data in the…

计算机视觉与模式识别 · 计算机科学 2023-12-11 Junjie Hu , Chenyou Fan , Mete Ozay , Hualie Jiang , Tin Lun Lam

Reinforcement learning (RL) struggles to scale to large, combinatorial action spaces common in many real-world problems. This paper introduces a novel framework for training discrete diffusion models as highly effective policies in these…

机器学习 · 计算机科学 2026-05-21 Haitong Ma , Ofir Nabati , Aviv Rosenberg , Bo Dai , Oran Lang , Craig Boutilier , Na Li , Shie Mannor , Lior Shani , Guy Tenneholtz

RGBD images, combining high-resolution color and lower-resolution depth from various types of depth sensors, are increasingly common. One can significantly improve the resolution of depth maps by taking advantage of color information; deep…

计算机视觉与模式识别 · 计算机科学 2019-09-10 Oleg Voynov , Alexey Artemov , Vage Egiazarian , Alexander Notchenko , Gleb Bobrovskikh , Denis Zorin , Evgeny Burnaev

While self-supervised monocular depth estimation in driving scenarios has achieved comparable performance to supervised approaches, violations of the static world assumption can still lead to erroneous depth predictions of traffic…

计算机视觉与模式识别 · 计算机科学 2023-08-22 Stefano Gasperini , Patrick Koch , Vinzenz Dallabetta , Nassir Navab , Benjamin Busam , Federico Tombari

RGB-D object tracking has attracted considerable attention recently, achieving promising performance thanks to the symbiosis between visual and depth channels. However, given a limited amount of annotated RGB-D tracking data, most…

计算机视觉与模式识别 · 计算机科学 2023-01-03 Xue-Feng Zhu , Tianyang Xu , Zhangyong Tang , Zucheng Wu , Haodong Liu , Xiao Yang , Xiao-Jun Wu , Josef Kittler

Face recognition in complex scenes suffers severe challenges coming from perturbations such as pose deformation, ill illumination, partial occlusion. Some methods utilize depth estimation to obtain depth corresponding to RGB to improve the…

计算机视觉与模式识别 · 计算机科学 2023-05-09 Wenhao Hu

We design a multiscopic vision system that utilizes a low-cost monocular RGB camera to acquire accurate depth estimation. Unlike multi-view stereo with images captured at unconstrained camera poses, the proposed system controls the motion…

计算机视觉与模式识别 · 计算机科学 2021-08-21 Weihao Yuan , Rui Fan , Michael Yu Wang , Qifeng Chen

Robots operating in unstructured environments require a comprehensive understanding of their surroundings, necessitating geometric and semantic information from sensor data. Traditional RGB-D processing pipelines focus primarily on…

计算机视觉与模式识别 · 计算机科学 2025-04-24 Zhiwu Zheng , Lauren Mentzer , Berk Iskender , Michael Price , Colm Prendergast , Audren Cloitre

Correspondence estimation is one of the most widely researched and yet only partially solved area of computer vision with many applications in tracking, mapping, recognition of objects and environment. In this paper, we propose a novel way…

计算机视觉与模式识别 · 计算机科学 2020-04-16 Umashankar Deekshith , Nishit Gajjar , Max Schwarz , Sven Behnke

The introduction of consumer RGB-D scanners set off a major boost in 3D computer vision research. Yet, the precision of existing depth scanners is not accurate enough to recover fine details of a scanned object. While modern shading based…

计算机视觉与模式识别 · 计算机科学 2016-03-31 Roy Or - El , Rom Hershkovitz , Aaron Wetzler , Guy Rosman , Alfred M. Bruckstein , Ron Kimmel

Diffusion models have achieved remarkable success in video generation; however, the high computational cost of the denoising process remains a major bottleneck. Existing approaches have shown promise in reducing the number of diffusion…

计算机视觉与模式识别 · 计算机科学 2026-02-03 Xiao Liang , Yunzhu Zhang , Linchao Zhu

We consider the problem of human pose estimation. While much recent work has focused on the RGB domain, these techniques are inherently under-constrained since there can be many 3D configurations that explain the same 2D projection. To this…

计算机视觉与模式识别 · 计算机科学 2019-11-19 Ren Li , Changjiang Cai , Georgios Georgakis , Srikrishna Karanam , Terrence Chen , Ziyan Wu

Diffusion magnetic resonance imaging (dMRI) is a crucial non-invasive technique for exploring the microstructure of the living human brain. Traditional hand-crafted and model-based tissue microstructure reconstruction methods often require…

图像与视频处理 · 电气工程与系统科学 2025-02-26 Xinrui Ma , Jian Cheng , Wenxin Fan , Ruoyou Wu , Yongquan Ye , Shanshan Wang

With the popularity of monocular videos generated by video sharing and live broadcasting applications, reconstructing and editing dynamic scenes in stationary monocular cameras has become a special but anticipated technology. In contrast to…

计算机视觉与模式识别 · 计算机科学 2024-02-02 Weixing Xie , Xiao Dong , Yong Yang , Qiqin Lin , Jingze Chen , Junfeng Yao , Xiaohu Guo

Applying data-driven approaches to non-rigid 3D reconstruction has been difficult, which we believe can be attributed to the lack of a large-scale training corpus. Unfortunately, this method fails for important cases such as highly…

计算机视觉与模式识别 · 计算机科学 2020-06-23 Aljaž Božič , Michael Zollhöfer , Christian Theobalt , Matthias Nießner

Accurate monocular depth estimation is crucial for 3D scene understanding, but existing methods often blur depth at object boundaries, introducing spurious intermediate 3D points. While achieving sharp edges usually requires very…

计算机视觉与模式识别 · 计算机科学 2025-11-19 Aurélien Cecille , Stefan Duffner , Franck Davoine , Rémi Agier , Thibault Neveu