English
Related papers

Related papers: MuPNet: Multi-modal Predictive Coding Network for …

200 papers

Robot navigation in unstructured environments requires multimodal perception systems that can support safe navigation. Multimodality enables the integration of complementary information collected by different sensors. However, this…

The advancement of deep learning has driven notable progress in remote sensing semantic segmentation. Attention mechanisms, while enabling global modeling and utilizing contextual information, face challenges of high computational costs and…

Computer Vision and Pattern Recognition · Computer Science 2024-04-23 Yang Yang , Shunyi Zheng

The recent vision transformer(i.e.for image classification) learns non-local attentive interaction of different patch tokens. However, prior arts miss learning the cross-scale dependencies of different pixels, the semantic correspondence of…

Computer Vision and Pattern Recognition · Computer Science 2021-06-29 Yuanfeng Ji , Ruimao Zhang , Huijie Wang , Zhen Li , Lingyun Wu , Shaoting Zhang , Ping Luo

This work proposes a perception system for autonomous vehicles and advanced driver assistance specialized on unpaved roads and off-road environments. In this research, the authors have investigated the behavior of Deep Learning algorithms…

In multi-person pose estimation, the left/right joint type discrimination is always a hard problem because of the similar appearance. Traditionally, we solve this problem by stacking multiple refinement modules to increase network's…

Computer Vision and Pattern Recognition · Computer Science 2019-11-27 Ying Huang , Jiankai Zhuang , Zengchang Qin

In response to the worldwide COVID-19 pandemic, advanced automated technologies have emerged as valuable tools to aid healthcare professionals in managing an increased workload by improving radiology report generation and prognostic…

Image and Video Processing · Electrical Eng. & Systems 2024-05-24 Zhusi Zhong , Jie Li , John Sollee , Scott Collins , Harrison Bai , Paul Zhang , Terrence Healey , Michael Atalay , Xinbo Gao , Zhicheng Jiao

Networks are ubiquitous structure that describes complex relationships between different entities in the real world. As a critical component of prediction task over nodes in networks, learning the feature representation of nodes has become…

Machine Learning · Computer Science 2018-09-10 Hansheng Xue , Jiajie Peng , Xuequn Shang

The area of Video Camouflaged Object Detection (VCOD) presents unique challenges in the field of computer vision due to texture similarities between target objects and their surroundings, as well as irregular motion patterns caused by both…

Computer Vision and Pattern Recognition · Computer Science 2024-02-05 Zifan Yu , Erfan Bank Tavakoli , Meida Chen , Suya You , Raghuveer Rao , Sanjeev Agarwal , Fengbo Ren

The ability to decompose scenes in terms of abstract building blocks is crucial for general intelligence. Where those basic building blocks share meaningful properties, interactions and other regularities across scenes, such decompositions…

Computer Vision and Pattern Recognition · Computer Science 2019-02-01 Christopher P. Burgess , Loic Matthey , Nicholas Watters , Rishabh Kabra , Irina Higgins , Matt Botvinick , Alexander Lerchner

The presence of task constraints imposes a significant challenge to motion planning. Despite all recent advancements, existing algorithms are still computationally expensive for most planning problems. In this paper, we present Constrained…

Robotics · Computer Science 2020-08-11 Ahmed H. Qureshi , Jiangeng Dong , Austin Choe , Michael C. Yip

Grounded Multimodal Named Entity Recognition (GMNER) is an emerging information extraction (IE) task, aiming to simultaneously extract entity spans, types, and corresponding visual regions of entities from given sentence-image pairs data.…

Information Retrieval · Computer Science 2025-01-28 Jielong Tang , Zhenxing Wang , Ziyang Gong , Jianxing Yu , Xiangwei Zhu , Jian Yin

Learning to reliably perceive and understand the scene is an integral enabler for robots to operate in the real-world. This problem is inherently challenging due to the multitude of object types as well as appearance changes caused by…

Computer Vision and Pattern Recognition · Computer Science 2021-11-05 Abhinav Valada , Rohit Mohan , Wolfram Burgard

We introduce MoNet, a novel functionally modular network for self-supervised and interpretable end-to-end learning. By leveraging its functional modularity with a latent-guided contrastive loss function, MoNet efficiently learns…

Machine Learning · Computer Science 2024-06-06 Hyunki Seong , David Hyunchul Shim

Image classification models often demonstrate unstable performance in real-world applications due to variations in image information, driven by differing visual perspectives of subject objects and lighting discrepancies. To mitigate these…

Computer Vision and Pattern Recognition · Computer Science 2024-07-29 Yuze Zheng , Zixuan Li , Xiangxian Li , Jinxing Liu , Yuqing Wang , Xiangxu Meng , Lei Meng

This work explores the use of spatial context as a source of free and plentiful supervisory signal for training a rich visual representation. Given only a large, unlabeled image collection, we extract random pairs of patches from each image…

Computer Vision and Pattern Recognition · Computer Science 2016-01-19 Carl Doersch , Abhinav Gupta , Alexei A. Efros

The place recognition problem comprises two distinct subproblems; recognizing a specific location in the world ("specific" or "ordinary" place recognition) and recognizing the type of place (place categorization). Both are important…

Robotics · Computer Science 2018-04-17 Sourav Garg , Adam Jacobson , Swagat Kumar , Michael Milford

Multimodal learning has shown significant performance boost compared to ordinary unimodal models across various domains. However, in real-world scenarios, multimodal signals are susceptible to missing because of sensor failures and adverse…

Computer Vision and Pattern Recognition · Computer Science 2025-09-16 Nhi Kieu , Kien Nguyen , Arnold Wiliem , Clinton Fookes , Sridha Sridharan

The task of semi-supervised video object segmentation (VOS) has been greatly advanced and state-of-the-art performance has been made by dense matching-based methods. The recent methods leverage space-time memory (STM) networks and learn to…

Computer Vision and Pattern Recognition · Computer Science 2021-11-30 Jiadai Sun , Yuxin Mao , Yuchao Dai , Yiran Zhong , Jianyuan Wang

Heterogeneous data fusion can enhance the robustness and accuracy of an algorithm on a given task. However, due to the difference in various modalities, aligning the sensors and embedding their information into discriminative and compact…

Computer Vision and Pattern Recognition · Computer Science 2022-11-01 Aditya Dutt , Alina Zare , Paul Gader

Semantic understanding and localization are fundamental enablers of robot autonomy that have for the most part been tackled as disjoint problems. While deep learning has enabled recent breakthroughs across a wide spectrum of scene…

Robotics · Computer Science 2018-10-12 Noha Radwan , Abhinav Valada , Wolfram Burgard