English
Related papers

Related papers: Efficient Multi-Task Scene Analysis with RGB-D Tra…

200 papers

Attention-based encoder-decoder framework is widely used in the scene text recognition task. However, for the current state-of-the-art(SOTA) methods, there is room for improvement in terms of the efficient usage of local visual and global…

Computer Vision and Pattern Recognition · Computer Science 2021-11-16 Mengmeng Cui , Wei Wang , Jinjin Zhang , Liang Wang

Vision transformer based models bring significant improvements for image segmentation tasks. Although these architectures offer powerful capabilities irrespective of specific segmentation tasks, their use of computational resources can be…

Computer Vision and Pattern Recognition · Computer Science 2026-04-01 Manyi Yao , Abhishek Aich , Yumin Suh , Amit Roy-Chowdhury , Christian Shelton , Manmohan Chandraker

Scene recognition is one of the basic problems in computer vision research with extensive applications in robotics. When available, depth images provide helpful geometric cues that complement the RGB texture information and help to identify…

Computer Vision and Pattern Recognition · Computer Science 2021-09-08 Andrea Ferreri , Silvia Bucci , Tatiana Tommasi

Recent advances in scene understanding benefit a lot from depth maps because of the 3D geometry information, especially in complex conditions (e.g., low light and overexposed). Existing approaches encode depth maps along with RGB images and…

Computer Vision and Pattern Recognition · Computer Science 2025-04-08 Bo-Wen Yin , Jiao-Long Cao , Ming-Ming Cheng , Qibin Hou

Indoor monocular semantic scene completion (MSSC) is notably more challenging than its outdoor counterpart due to complex spatial layouts and severe occlusions. While transformers are well suited for modeling global dependencies, their high…

Computer Vision and Pattern Recognition · Computer Science 2026-03-27 Xuzhi Wang , Xinran Wu , Song Wang , Lingdong Kong , Ziping Zhao

Recently, Transformer-based methods have achieved impressive results in single image super-resolution (SISR). However, the lack of locality mechanism and high complexity limit their application in the field of super-resolution (SR). To…

Computer Vision and Pattern Recognition · Computer Science 2023-06-21 Ling Zheng , Jinchen Zhu , Jinpeng Shi , Shizhuang Weng

Scene understanding is paramount in robotics, self-navigation, augmented reality, and many other fields. To fully accomplish this task, an autonomous agent has to infer the 3D structure of the sensed scene (to know where it looks at) and…

Computer Vision and Pattern Recognition · Computer Science 2020-02-26 Pier Luigi Dovesi , Matteo Poggi , Lorenzo Andraghetti , Miquel Martí , Hedvig Kjellström , Alessandro Pieropan , Stefano Mattoccia

Pattern recognition based on RGB-Event data is a newly arising research topic and previous works usually learn their features using CNN or Transformer. As we know, CNN captures the local features well and the cascaded self-attention…

Computer Vision and Pattern Recognition · Computer Science 2025-04-21 Xiao Wang , Yao Rong , Shiao Wang , Yuan Chen , Zhe Wu , Bo Jiang , Yonghong Tian , Jin Tang

High-speed vision sensing is essential for real-time perception in applications such as robotics, autonomous vehicles, and industrial automation. Traditional frame-based vision systems suffer from motion blur, high latency, and redundant…

Computer Vision and Pattern Recognition · Computer Science 2025-07-10 Riadul Islam , Joey Mulé , Dhandeep Challagundla , Shahmir Rizvi , Sean Carson

Robots operating in unstructured environments require a comprehensive understanding of their surroundings, necessitating geometric and semantic information from sensor data. Traditional RGB-D processing pipelines focus primarily on…

Computer Vision and Pattern Recognition · Computer Science 2025-04-24 Zhiwu Zheng , Lauren Mentzer , Berk Iskender , Michael Price , Colm Prendergast , Audren Cloitre

A long-standing goal in scene understanding is to obtain interpretable and editable representations that can be directly constructed from a raw monocular RGB-D video, without requiring specialized hardware setup or priors. The problem is…

Computer Vision and Pattern Recognition · Computer Science 2023-06-22 Yu-Shiang Wong , Niloy J. Mitra

Recently, deep learning methods have achieved state-of-the-art performance in many medical image segmentation tasks. Many of these are based on convolutional neural networks (CNNs). For such methods, the encoder is the key part for global…

Image and Video Processing · Electrical Eng. & Systems 2022-08-25 Hao Li , Dewei Hu , Han Liu , Jiacheng Wang , Ipek Oguz

RGB-D has gradually become a crucial data source for understanding complex scenes in assisted driving. However, existing studies have paid insufficient attention to the intrinsic spatial properties of depth maps. This oversight…

Computer Vision and Pattern Recognition · Computer Science 2025-07-02 Siyu Chen , Ting Han , Changshe Zhang , Weiquan Liu , Jinhe Su , Zongyue Wang , Guorong Cai

Decision-making for urban autonomous driving is challenging due to the stochastic nature of interactive traffic participants and the complexity of road structures. Although reinforcement learning (RL)-based decision-making scheme is…

Machine Learning · Computer Science 2023-08-28 Haochen Liu , Zhiyu Huang , Xiaoyu Mo , Chen Lv

We introduce TransformerFusion, a transformer-based 3D scene reconstruction approach. From an input monocular RGB video, the video frames are processed by a transformer network that fuses the observations into a volumetric feature grid…

Computer Vision and Pattern Recognition · Computer Science 2021-07-07 Aljaž Božič , Pablo Palafox , Justus Thies , Angela Dai , Matthias Nießner

Deep convolutional networks (CNN) can achieve impressive results on RGB scene recognition thanks to large datasets such as Places. In contrast, RGB-D scene recognition is still underdeveloped in comparison, due to two limitations of RGB-D…

Computer Vision and Pattern Recognition · Computer Science 2018-10-30 Xinhang Song , Shuqiang Jiang , Luis Herranz , Chengpeng Chen

The recent advancements in deep convolutional neural networks have shown significant promise in the domain of road scene parsing. Nevertheless, the existing works focus primarily on freespace detection, with little attention given to…

Computer Vision and Pattern Recognition · Computer Science 2024-07-02 Jiahang Li , Yikang Zhang , Peng Yun , Guangliang Zhou , Qijun Chen , Rui Fan

Scene understanding is an active research area. Commercial depth sensors, such as Kinect, have enabled the release of several RGB-D datasets over the past few years which spawned novel methods in 3D scene understanding. More recently with…

Computer Vision and Pattern Recognition · Computer Science 2022-01-13 Gilad Baruch , Zhuoyuan Chen , Afshin Dehghan , Tal Dimry , Yuri Feigin , Peter Fu , Thomas Gebauer , Brandon Joffe , Daniel Kurz , Arik Schwartz , Elad Shulman

Density map estimation enables accurate object counting in heavily occluded, and densely packed scenes where detection-based counting fails. In multi-class density estimation, class awareness can be introduced by modelling classes…

Computer Vision and Pattern Recognition · Computer Science 2026-04-16 Villanelle O'Reilly , Jonathan Cox , Georgios Leontidis , Marc Hanheide , Petra Bosilj , James M. Brown

Inferring walls configuration of indoor environment could help robot "understand" the environment better. This allows the robot to execute a task that involves inter-room navigation, such as picking an object in the kitchen. In this paper,…

Robotics · Computer Science 2018-03-29 Ismail Rusli , Bambang Riyanto Trilaksono , Widyawardana Adiprawita