English
Related papers

Related papers: Cross-Modal Attentional Context Learning for RGB-D…

200 papers

The deployment of machine learning models in safety-critical applications comes with the expectation that such models will perform well over a range of contexts (e.g., a vision model for classifying street signs should work in rural, city,…

Computer Vision and Pattern Recognition · Computer Science 2022-11-29 Nathan Drenkow , Alvin Tan , Chace Ashcraft , Kiran Karra

Multimodal object detection has attracted significant attention in both academia and industry for its enhanced robustness. Although numerous studies have focused on improving modality fusion strategies, most neglect fusion degradation, and…

Computer Vision and Pattern Recognition · Computer Science 2025-11-20 YiKang Shao , Tao Shi

Contextual information, such as the co-occurrence of objects and the spatial and relative size among objects provides deep and complex information about scenes. It also can play an important role in improving object detection. In this work,…

Computer Vision and Pattern Recognition · Computer Science 2019-06-07 Faisal Alamri , Nicolas Pugeault

Change detection encompasses a variety of task types, and the goal of building change detection (BCD) tasks is to accurately locate buildings and distinguish changed building areas. In recent years, various deep learning-based BCD methods…

Image and Video Processing · Electrical Eng. & Systems 2026-03-11 ChengMing Wang

The proliferation of sophisticated AI-generated deepfakes poses critical challenges for digital media authentication and societal security. While existing detection methods perform well within specific generative domains, they exhibit…

Computer Vision and Pattern Recognition · Computer Science 2025-05-26 Naseem Khan , Tuan Nguyen , Amine Bermak , Issa Khalil

Multi-modal 3D object understanding has gained significant attention, yet current approaches often assume complete data availability and rigid alignment across all modalities. We present CrossOver, a novel framework for cross-modal 3D scene…

Computer Vision and Pattern Recognition · Computer Science 2025-04-08 Sayan Deb Sarkar , Ondrej Miksik , Marc Pollefeys , Daniel Barath , Iro Armeni

Detecting hidden or partially concealed objects remains a fundamental challenge in multimodal environments, where factors like occlusion, camouflage, and lighting variations significantly hinder performance. Traditional RGB-based detection…

Computer Vision and Pattern Recognition · Computer Science 2025-09-15 Harris Song , Tuan-Anh Vu , Sanjith Menon , Sriram Narasimhan , M. Khalid Jawed

Scene recognition with RGB images has been extensively studied and has reached very remarkable recognition levels, thanks to convolutional neural networks (CNN) and large scene datasets. In contrast, current RGB-D scene data is much more…

Computer Vision and Pattern Recognition · Computer Science 2018-01-23 Xinhang Song , Luis Herranz , Shuqiang Jiang

Driver action recognition has significantly advanced in enhancing driver-vehicle interactions and ensuring driving safety by integrating multiple modalities, such as infrared and depth. Nevertheless, compared to RGB modality only, it is…

Computer Vision and Pattern Recognition · Computer Science 2024-08-06 Ruoyu Wang , Chen Cai , Wenqian Wang , Jianjun Gao , Dan Lin , Wenyang Liu , Kim-Hui Yap

Recognizing objects in images is a fundamental problem in computer vision. Although detecting objects in 2D images is common, many applications require determining their pose in 3D space. Traditional category-level methods rely on RGB-D…

Computer Vision and Pattern Recognition · Computer Science 2025-08-05 Tom Fischer , Xiaojie Zhang , Eddy Ilg

Much of the focus in the object detection literature has been on the problem of identifying the bounding box of a particular class of object in an image. Yet, in contexts such as robotics and augmented reality, it is often necessary to find…

Computer Vision and Pattern Recognition · Computer Science 2020-11-17 Jean-Philippe Mercier , Mathieu Garon , Philippe Giguère , Jean-François Lalonde

Crowd counting research has made significant advancements in real-world applications, but it remains a formidable challenge in cross-modal settings. Most existing methods rely solely on the optical features of RGB images, ignoring the…

Computer Vision and Pattern Recognition · Computer Science 2022-11-15 Youjia Zhang , Soyun Choi , Sungeun Hong

In this paper, we address the 3D object detection task by capturing multi-level contextual information with the self-attention mechanism and multi-scale feature fusion. Most existing 3D object detection methods recognize objects…

Computer Vision and Pattern Recognition · Computer Science 2020-04-14 Qian Xie , Yu-Kun Lai , Jing Wu , Zhoutao Wang , Yiming Zhang , Kai Xu , Jun Wang

In this paper, we propose a novel method for plane clustering specialized in cluttered scenes using an RGB-D camera and validate its effectiveness through robot grasping experiments. Unlike existing methods, which focus on large-scale…

Robotics · Computer Science 2024-03-20 Seunghyeon Lim , Youngjae Yoo , Jun Ki Lee , Byoung-Tak Zhang

We study the problem of object detection over scanned images of scientific documents. We consider images that contain objects of varying aspect ratios and sizes and range from coarse elements such as tables and figures to fine elements such…

Computer Vision and Pattern Recognition · Computer Science 2019-10-31 Ankur Goswami , Joshua McGrath , Shanan Peters , Theodoros Rekatsinas

In recent years, object detection utilizing both visible (RGB) and thermal infrared (IR) imagery has garnered extensive attention and has been widely implemented across a diverse array of fields. By leveraging the complementary properties…

Computer Vision and Pattern Recognition · Computer Science 2025-12-30 Tianyi Zhao , Maoxun Yuan , Feng Jiang , Nan Wang , Xingxing Wei

Encoder-decoder models have been widely used in RGBD semantic segmentation, and most of them are designed via a two-stream network. In general, jointly reasoning the color and geometric information from RGBD is beneficial for semantic…

Computer Vision and Pattern Recognition · Computer Science 2022-03-16 Yang Zhang , Yang Yang , Chenyun Xiong , Guodong Sun , Yanwen Guo

Multi-modal industrial anomaly detection typically relies on separate models for each product category, fundamentally limiting practical scalability. When shifting to a unified paradigm that handles diverse classes simultaneously, detection…

Computer Vision and Pattern Recognition · Computer Science 2026-05-29 Yangchen Wu , Huiqiang Xie

Connectionist Temporal Classification (CTC) and attention mechanism are two main approaches used in recent scene text recognition works. Compared with attention-based methods, CTC decoder has a much shorter inference time, yet a lower…

Computer Vision and Pattern Recognition · Computer Science 2020-02-05 Wenyang Hu , Xiaocong Cai , Jun Hou , Shuai Yi , Zhiping Lin

Context, as referred to situational factors related to the object of interest, can help infer the object's states or properties in visual recognition. As such contextual features are too diverse (across instances) to be annotated, existing…

Computer Vision and Pattern Recognition · Computer Science 2021-10-11 Mingzhou Liu , Xinwei Sun , Fandong Zhang , Yizhou Yu , Yizhou Wang
‹ Prev 1 8 9 10 Next ›