中文
相关论文

相关论文: PDDM: Pseudo Depth Diffusion Model for RGB-PD Sema…

200 篇论文

Previous RGB-D salient object detection (SOD) methods have widely adopted deep learning tools to automatically strike a trade-off between RGB and D (depth), whose key rationale is to take full advantage of their complementary nature, aiming…

计算机视觉与模式识别 · 计算机科学 2020-08-11 Xuehao Wang , Shuai Li , Chenglizhao Chen , Aimin Hao , Hong Qin

Depth maps produced by consumer-grade sensors suffer from inaccurate measurements and missing data from either system or scene-specific sources. Data-driven denoising algorithms can mitigate such problems. However, they require vast amounts…

计算机视觉与模式识别 · 计算机科学 2024-07-04 Alexandre Duarte , Francisco Fernandes , João M. Pereira , Catarina Moreira , Jacinto C. Nascimento , Joaquim Jorge

Accurate three-dimensional perception is a fundamental task in several computer vision applications. Recently, commercial RGB-depth (RGB-D) cameras have been widely adopted as single-view depth-sensing devices owing to their efficient…

计算机视觉与模式识别 · 计算机科学 2022-07-27 Jiwan Kim , Minchang Kim , Yeong-Gil Shin , Minyoung Chung

This work introduces RGBX-DiffusionDet, an object detection framework extending the DiffusionDet model to fuse the heterogeneous 2D data (X) with RGB imagery via an adaptive multimodal encoder. To enable cross-modal interaction, we design…

计算机视觉与模式识别 · 计算机科学 2026-01-05 Eliraz Orfaig , Inna Stainvas , Igal Bilik

Depth estimation in complex real-world scenarios is a challenging task, especially when relying solely on a single modality such as visible light or thermal infrared (THR) imagery. This paper proposes a novel multimodal depth estimation…

图像与视频处理 · 电气工程与系统科学 2025-04-30 Zelin Meng , Takanori Fukao

RGB-based semantic segmentation has become a mainstream approach for visual perception and is widely applied in a variety of downstream tasks. However, existing methods typically rely on high-resolution RGB inputs, which may expose…

机器人学 · 计算机科学 2026-04-07 Xuying Huang , Sicong Pan , Olga Zatsarynna , Juergen Gall , Maren Bennewitz

In the realm of computer vision, the integration of advanced techniques into the processing of RGB-D camera inputs poses a significant challenge, given the inherent complexities arising from diverse environmental conditions and varying…

计算机视觉与模式识别 · 计算机科学 2024-05-02 Safouane El Ghazouali , Youssef Mhirit , Ali Oukhrid , Umberto Michelucci , Hichem Nouira

In this paper, we introduce Semi-SMD, a novel metric depth estimation framework tailored for surrounding cameras equipment in autonomous driving. In this work, the input data consists of adjacent surrounding frames and camera parameters. We…

机器人学 · 计算机科学 2025-09-10 Yusen Xie , Zhengmin Huang , Shaojie Shen , Jun Ma

Current deep learning approaches in computer vision primarily focus on RGB data sacrificing information. In contrast, RAW images offer richer representation, which is crucial for precise recognition, particularly in challenging conditions…

计算机视觉与模式识别 · 计算机科学 2024-11-21 Christoph Reinders , Radu Berdan , Beril Besbinar , Junji Otsuka , Daisuke Iso

Consumer-level depth cameras and depth sensors embedded in mobile devices enable numerous applications, such as AR games and face identification. However, the quality of the captured depth is sometimes insufficient for 3D reconstruction,…

Existing RGB-D saliency detection models do not explicitly encourage RGB and depth to achieve effective multi-modal learning. In this paper, we introduce a novel multi-stage cascaded learning framework via mutual information minimization to…

计算机视觉与模式识别 · 计算机科学 2022-01-07 Jing Zhang , Deng-Ping Fan , Yuchao Dai , Xin Yu , Yiran Zhong , Nick Barnes , Ling Shao

Differentially private SGD (DP-SGD) is one of the most popular methods for solving differentially private empirical risk minimization (ERM). Due to its noisy perturbation on each gradient update, the error rate of DP-SGD scales with the…

机器学习 · 计算机科学 2021-04-27 Yingxue Zhou , Zhiwei Steven Wu , Arindam Banerjee

Latent diffusion models have proven to be state-of-the-art in the creation and manipulation of visual outputs. However, as far as we know, the generation of depth maps jointly with RGB is still limited. We introduce LDM3D-VR, a suite of…

计算机视觉与模式识别 · 计算机科学 2023-11-07 Gabriela Ben Melech Stan , Diana Wofk , Estelle Aflalo , Shao-Yen Tseng , Zhipeng Cai , Michael Paulitsch , Vasudev Lal

Most of the existing visual SLAM methods heavily rely on a static world assumption and easily fail in dynamic environments. Some recent works eliminate the influence of dynamic objects by introducing deep learning-based semantic information…

机器人学 · 计算机科学 2022-01-10 Tete Ji , Chen Wang , Lihua Xie

Given the widespread adoption of depth-sensing acquisition devices, RGB-D videos and related data/media have gained considerable traction in various aspects of daily life. Consequently, conducting salient object detection (SOD) in RGB-D…

计算机视觉与模式识别 · 计算机科学 2024-05-22 Ao Mou , Yukang Lu , Jiahao He , Dingyao Min , Keren Fu , Qijun Zhao

This work addresses multi-class segmentation of indoor scenes with RGB-D inputs. While this area of research has gained much attention recently, most works still rely on hand-crafted features. In contrast, we apply a multiscale…

计算机视觉与模式识别 · 计算机科学 2013-03-15 Camille Couprie , Clément Farabet , Laurent Najman , Yann LeCun

Multi-modal depth estimation is one of the key challenges for endowing autonomous machines with robust robotic perception capabilities. There have been outstanding advances in the development of uni-modal depth estimation techniques based…

机器人学 · 计算机科学 2023-07-21 Johan S. Obando-Ceron , Victor Romero-Cano , Sildomar Monteiro

Most existing RGB-D semantic segmentation methods focus on the feature level fusion, including complex cross-modality and cross-scale fusion modules. However, these methods may cause misalignment problem in the feature fusion process and…

计算机视觉与模式识别 · 计算机科学 2025-05-05 Xiaoyan Jiang , Bohan Wang , Xinlong Wan , Shanshan Chen , Hamido Fujita , Hanan Abd. Al Juaid

Multi-sensor fusion has significant potential in perception tasks for both indoor and outdoor environments. Especially under challenging conditions such as adverse weather and low-light environments, the combined use of millimeter-wave…

图像与视频处理 · 电气工程与系统科学 2025-05-23 Tieshuai Song , Jiandong Ye , Ao Guo , Guidong He , Bin Yang

This paper presents a GPU implementation of two foreground object segmentation algorithms: Gaussian Mixture Model (GMM) and Pixel Based Adaptive Segmenter (PBAS) modified for RGB-D data support. The simultaneous use of colour (RGB) and…

计算机视觉与模式识别 · 计算机科学 2020-07-02 Piotr Janus , Tomasz Kryjak , Marek Gorgon