English
Related papers

Related papers: Depth Matters: Multimodal RGB-D Perception for Rob…

200 papers

Depth sensing is crucial for 3D reconstruction and scene understanding. Active depth sensors provide dense metric measurements, but often suffer from limitations such as restricted operating ranges, low spatial resolution, sensor…

Computer Vision and Pattern Recognition · Computer Science 2019-01-10 Chao Liu , Jinwei Gu , Kihwan Kim , Srinivasa Narasimhan , Jan Kautz

The focal point of egocentric video understanding is modelling hand-object interactions. Standard models -- CNNs, Vision Transformers, etc. -- which receive RGB frames as input perform well, however, their performance improves further by…

Computer Vision and Pattern Recognition · Computer Science 2022-10-11 Gorjan Radevski , Dusan Grujicic , Matthew Blaschko , Marie-Francine Moens , Tinne Tuytelaars

Most of the state-of-the-art indirect visual SLAM methods are based on the sparse point features. However, it is hard to find enough reliable point features for state estimation in the case of low-textured scenes. Line features are abundant…

Robotics · Computer Science 2021-02-16 Xin Ma , Xinwu Liang

In robotic vision, a de-facto paradigm is to learn in simulated environments and then transfer to real-world applications, which poses an essential challenge in bridging the sim-to-real domain gap. While mainstream works tackle this problem…

Computer Vision and Pattern Recognition · Computer Science 2024-04-08 Xingyu Liu , Chenyangguang Zhang , Gu Wang , Ruida Zhang , Xiangyang Ji

RGBD images, combining high-resolution color and lower-resolution depth from various types of depth sensors, are increasingly common. One can significantly improve the resolution of depth maps by taking advantage of color information; deep…

Computer Vision and Pattern Recognition · Computer Science 2019-09-10 Oleg Voynov , Alexey Artemov , Vage Egiazarian , Alexander Notchenko , Gleb Bobrovskikh , Denis Zorin , Evgeny Burnaev

Multi-sensor fusion has significant potential in perception tasks for both indoor and outdoor environments. Especially under challenging conditions such as adverse weather and low-light environments, the combined use of millimeter-wave…

Image and Video Processing · Electrical Eng. & Systems 2025-05-23 Tieshuai Song , Jiandong Ye , Ao Guo , Guidong He , Bin Yang

Human volumetric capture is a long-standing topic in computer vision and computer graphics. Although high-quality results can be achieved using sophisticated off-line systems, real-time human volumetric capture of complex scenarios,…

Computer Vision and Pattern Recognition · Computer Science 2021-05-07 Tao Yu , Zerong Zheng , Kaiwen Guo , Pengpeng Liu , Qionghai Dai , Yebin Liu

In this paper, we propose a new correlated and individual multi-modal deep learning (CIMDL) method for RGB-D object recognition. Unlike most conventional RGB-D object recognition methods which extract features from the RGB and depth…

Computer Vision and Pattern Recognition · Computer Science 2016-12-12 Ziyan Wang , Jiwen Lu , Ruogu Lin , Jianjiang Feng , Jie zhou

We present a control approach for autonomous vehicles based on deep reinforcement learning. A neural network agent is trained to map its estimated state to acceleration and steering commands given the objective of reaching a specific target…

Robotics · Computer Science 2020-03-16 Andreas Folkers , Matthias Rick , Christof Büskens

Collaborative visual perception methods have gained widespread attention in the autonomous driving community in recent years due to their ability to address sensor limitation problems. However, the absence of explicit depth information…

Computer Vision and Pattern Recognition · Computer Science 2025-10-10 Shaohong Wang , Bin Lu , Xinyu Xiao , Hanzhi Zhong , Bowen Pang , Tong Wang , Zhiyu Xiang , Hangguan Shan , Eryun Liu

Model-based reinforcement learning (MBRL) techniques have recently yielded promising results for real-world autonomous racing using high-dimensional observations. MBRL agents, such as Dreamer, solve long-horizon tasks by building a world…

Robotics · Computer Science 2023-05-09 Elena Shrestha , Chetan Reddy , Hanxi Wan , Yulun Zhuang , Ram Vasudevan

Visual navigation for autonomous agents is a core task in the fields of computer vision and robotics. Learning-based methods, such as deep reinforcement learning, have the potential to outperform the classical solutions developed for this…

Computer Vision and Pattern Recognition · Computer Science 2021-03-23 Zachary Seymour , Kowshik Thopalli , Niluthpol Mithun , Han-Pang Chiu , Supun Samarasekera , Rakesh Kumar

Cooperative perception enhances autonomous driving by leveraging Vehicle-to-Everything (V2X) communication for multi-agent sensor fusion. However, most existing methods rely on single-modal data sharing, limiting fusion performance,…

Robotics · Computer Science 2025-09-25 Lantao Li , Kang Yang , Wenqi Zhang , Xiaoxue Wang , Chen Sun

Recently, it is increasingly popular to equip mobile RGB cameras with Time-of-Flight (ToF) sensors for active depth sensing. However, for off-the-shelf ToF sensors, one must tackle two problems in order to obtain high-quality depth with…

Computer Vision and Pattern Recognition · Computer Science 2019-09-18 Di Qiu , Jiahao Pang , Wenxiu Sun , Chengxi Yang

Existing RGB-D saliency detection models do not explicitly encourage RGB and depth to achieve effective multi-modal learning. In this paper, we introduce a novel multi-stage cascaded learning framework via mutual information minimization to…

Computer Vision and Pattern Recognition · Computer Science 2022-01-07 Jing Zhang , Deng-Ping Fan , Yuchao Dai , Xin Yu , Yiran Zhong , Nick Barnes , Ling Shao

Depth (D) indicates occlusion and is less sensitive to illumination changes, which make depth attractive modality for Visual Object Tracking (VOT). Depth is used in RGBD object tracking where the best trackers are deep RGB trackers with…

Computer Vision and Pattern Recognition · Computer Science 2021-10-25 Song Yan , Jinyu Yang , Ales Leonardis , Joni-Kristian Kamarainen

Deep learning approaches have achieved highly accurate face recognition by training the models with very large face image datasets. Unlike the availability of large 2D face image datasets, there is a lack of large 3D face datasets available…

Computer Vision and Pattern Recognition · Computer Science 2021-12-23 Meng-Tzu Chiu , Hsun-Ying Cheng , Chien-Yi Wang , Shang-Hong Lai

The reasonable employment of RGB and depth data show great significance in promoting the development of computer vision tasks and robot-environment interaction. However, there are different advantages and disadvantages in the early and late…

Computer Vision and Pattern Recognition · Computer Science 2021-09-13 Jinchao Zhu

This study aims to improve the performance and generalization capability of end-to-end autonomous driving with scene understanding leveraging deep learning and multimodal sensor fusion techniques. The designed end-to-end deep neural network…

Robotics · Computer Science 2020-08-04 Zhiyu Huang , Chen Lv , Yang Xing , Jingda Wu

We propose a cross attention transformer based method for multimodal sensor fusion to build a birds eye view of a vessels surroundings supporting safer autonomous marine navigation. The model deeply fuses multiview RGB and long wave…

Computer Vision and Pattern Recognition · Computer Science 2025-05-06 Dimitrios Dagdilelis , Panagiotis Grigoriadis , Roberto Galeazzi