English
Related papers

Related papers: DFormer: Rethinking RGBD Representation Learning f…

200 papers

The integration of RGB and depth modalities significantly enhances the accuracy of segmenting complex indoor scenes, with depth data from RGB-D cameras playing a crucial role in this improvement. However, collecting an RGB-D dataset is more…

Computer Vision and Pattern Recognition · Computer Science 2025-03-25 Xinhua Xu , Hong Liu , Jianbing Wu , Jinfu Liu

Multi-modal RGB and Depth (RGBD) data are predominant in many domains such as robotics, autonomous driving and remote sensing. The combination of these multi-modal data enhances environmental perception by providing 3D spatial context,…

Computer Vision and Pattern Recognition · Computer Science 2025-06-02 Roger Ferrod , Cássio F. Dantas , Luigi Di Caro , Dino Ienco

Crack segmentation is crucial in civil engineering, particularly for assessing pavement integrity and ensuring the durability of infrastructure. While deep learning has advanced RGB-based segmentation, performance degrades under adverse…

Computer Vision and Pattern Recognition · Computer Science 2024-12-31 Ruiqiang Xiao , Xiaohu Chen

Recently, heatmap regression methods based on 1D landmark representations have shown prominent performance on locating facial landmarks. However, previous methods ignored to make deep explorations on the good potentials of 1D landmark…

Computer Vision and Pattern Recognition · Computer Science 2024-02-02 Shi Yin , Shijie Huan , Shangfei Wang , Jinshui Hu , Tao Guo , Bing Yin , Baocai Yin , Cong Liu

Segmenting unseen objects in cluttered scenes is an important skill that robots need to acquire in order to perform tasks in new environments. In this work, we propose a new method for unseen object instance segmentation by learning RGB-D…

Robotics · Computer Science 2021-03-04 Yu Xiang , Christopher Xie , Arsalan Mousavian , Dieter Fox

The adoption of Vision Transformers (ViTs) based architectures represents a significant advancement in 3D Medical Image (MI) segmentation, surpassing traditional Convolutional Neural Network (CNN) models by enhancing global contextual…

Computer Vision and Pattern Recognition · Computer Science 2024-04-25 Shehan Perera , Pouyan Navard , Alper Yilmaz

Transformer has achieved great successes in learning vision and language representation, which is general across various downstream tasks. In visual control, learning transferable state representation that can transfer between different…

Computer Vision and Pattern Recognition · Computer Science 2022-06-20 Yao Mu , Shoufa Chen , Mingyu Ding , Jianyu Chen , Runjian Chen , Ping Luo

Existing Blind image Super-Resolution (BSR) methods focus on estimating either kernel or degradation information, but have long overlooked the essential content details. In this paper, we propose a novel BSR approach, Content-aware…

Computer Vision and Pattern Recognition · Computer Science 2024-07-02 Qingguo Liu , Chenyi Zhuang , Pan Gao , Jie Qin

Despite the tremendous progress in zero-shot learning(ZSL), the majority of existing methods still rely on human-annotated attributes, which are difficult to annotate and scale. An unsupervised alternative is to represent each class using…

Computer Vision and Pattern Recognition · Computer Science 2022-09-22 Muhammad Ferjad Naeem , Yongqin Xian , Luc Van Gool , Federico Tombari

We present a High-Resolution Transformer (HRFormer) that learns high-resolution representations for dense prediction tasks, in contrast to the original Vision Transformer that produces low-resolution representations and has high memory and…

Computer Vision and Pattern Recognition · Computer Science 2021-11-09 Yuhui Yuan , Rao Fu , Lang Huang , Weihong Lin , Chao Zhang , Xilin Chen , Jingdong Wang

This paper presents a novel deep neural network framework for RGB-D salient object detection by controlling the message passing between the RGB images and depth maps on the feature level and exploring the long-range semantic contexts and…

Computer Vision and Pattern Recognition · Computer Science 2022-06-22 Baian Chen , Zhilei Chen , Xiaowei Hu , Jun Xu , Haoran Xie , Mingqiang Wei , Jing Qin

Ground-truth RGBD data are fundamental for a wide range of computer vision applications; however, those labeled samples are difficult to collect and time-consuming to produce. A common solution to overcome this lack of data is to employ…

Computer Vision and Pattern Recognition · Computer Science 2024-05-28 L. Papa , P. Russo , I. Amerini

We introduce SceneNet RGB-D, expanding the previous work of SceneNet to enable large scale photorealistic rendering of indoor scene trajectories. It provides pixel-perfect ground truth for scene understanding problems such as semantic…

Computer Vision and Pattern Recognition · Computer Science 2017-01-31 John McCormac , Ankur Handa , Stefan Leutenegger , Andrew J. Davison

Semantic image and video segmentation stand among the most important tasks in computer vision nowadays, since they provide a complete and meaningful representation of the environment by means of a dense classification of the pixels in a…

Computer Vision and Pattern Recognition · Computer Science 2023-03-09 Felipe Manfio Barbosa , Fernando Santos Osório

Planar grasp detection is one of the most fundamental tasks to robotic manipulation, and the recent progress of consumer-grade RGB-D sensors enables delivering more comprehensive features from both the texture and shape modalities. However,…

Robotics · Computer Science 2023-03-01 Ran Qin , Haoxiang Ma , Boyang Gao , Di Huang

Recent work on visual representation learning has shown to be efficient for robotic manipulation tasks. However, most existing works pretrained the visual backbone solely on 2D images or egocentric videos, ignoring the fact that robots…

RGBD (RGB plus depth) object tracking is gaining momentum as RGBD sensors have become popular in many application fields such as robotics.However, the best RGBD trackers are extensions of the state-of-the-art deep RGB trackers. They are…

Computer Vision and Pattern Recognition · Computer Science 2021-09-01 Song Yan , Jinyu Yang , Jani Käpylä , Feng Zheng , Aleš Leonardis , Joni-Kristian Kämäräinen

Referring Remote Sensing Image Segmentation (RRSIS) is a challenging task, aiming to segment specific target objects in remote sensing (RS) images based on a given language expression. Existing RRSIS methods typically employ coarse-grained…

Computer Vision and Pattern Recognition · Computer Science 2025-04-01 Maofu Liu , Xin Jiang , Xiaokang Zhang

Semantic segmentation of remotely sensed urban scene images is required in a wide range of practical applications, such as land cover mapping, urban change detection, environmental protection, and economic assessment.Driven by rapid…

Computer Vision and Pattern Recognition · Computer Science 2022-06-28 Libo Wang , Rui Li , Ce Zhang , Shenghui Fang , Chenxi Duan , Xiaoliang Meng , Peter M. Atkinson

Vision-language models (VLMs) have been widely applied to 2D medical image analysis due to their ability to align visual and textual representations. However, extending VLMs to 3D imaging remains computationally challenging. Existing 3D…

Computer Vision and Pattern Recognition · Computer Science 2025-04-28 Gorkem Can Ates , Yu Xin , Kuang Gong , Wei Shao