English
Related papers

Related papers: UniMPR: A Unified Framework for Multimodal Place R…

200 papers

The growing availability of co-located geospatial data spanning aerial imagery, street-level views, elevation models, text, and geographic coordinates offers a unique opportunity for multimodal representation learning. We introduce…

Computer Vision and Pattern Recognition · Computer Science 2026-04-14 Guillaume Astruc , Eduard Trulls , Jan Hosang , Loic Landrieu , Paul-Edouard Sarlin

Current remote sensing change detection (CD) methods mainly rely on specialized models, which limits the scalability toward modality-adaptive Earth observation. For homogeneous CD, precise boundary delineation relies on fine-grained spatial…

Computer Vision and Pattern Recognition · Computer Science 2026-01-22 Qingling Shu , Sibao Chen , Wei Lu , Zhihui You , Chengzhuang Liu

Fusion-based place recognition is an emerging technique jointly utilizing multi-modal perception data, to recognize previously visited places in GPS-denied scenarios for robots and autonomous vehicles. Recent fusion-based place recognition…

Computer Vision and Pattern Recognition · Computer Science 2024-02-28 Jingyi Xu , Junyi Ma , Qi Wu , Zijie Zhou , Yue Wang , Xieyuanli Chen , Ling Pei

Predicting the behavior of road users accurately is crucial to enable the safe operation of autonomous vehicles in urban or densely populated areas. Therefore, there has been a growing interest in time series motion prediction research,…

Machine Learning · Computer Science 2024-10-22 Camiel Oerlemans , Bram Grooten , Michiel Braat , Alaa Alassi , Emilia Silvas , Decebal Constantin Mocanu

This article introduces BEVPlace++, a novel, fast, and robust LiDAR global localization method for unmanned ground vehicles. It uses lightweight convolutional neural networks (CNNs) on Bird's Eye View (BEV) image-like representations of…

Robotics · Computer Science 2025-06-26 Lun Luo , Si-Yuan Cao , Xiaorui Li , Jintao Xu , Rui Ai , Zhu Yu , Xieyuanli Chen

Recent advances in Large Multi-modal Models (LMMs) have demonstrated their remarkable success as general-purpose multi-modal assistants, with particular focuses on holistic image- and video-language understanding. Conversely, less attention…

Computer Vision and Pattern Recognition · Computer Science 2025-11-11 Ye Liu , Zongyang Ma , Junfu Pu , Zhongang Qi , Yang Wu , Ying Shan , Chang Wen Chen

A unified representation space in multi-modal learning is essential for effectively integrating diverse data sources, such as text, images, and audio, to enhance efficiency and performance across various downstream tasks. Recent binding…

Machine Learning · Computer Science 2025-10-08 Minoh Jeong , Zae Myung Kim , Min Namgung , Dongyeop Kang , Yao-Yi Chiang , Alfred Hero

Multimodal representation learning techniques typically rely on paired samples to learn common representations, but paired samples are challenging to collect in fields such as biology where measurement devices often destroy the samples.…

Machine Learning · Computer Science 2024-10-30 Johnny Xi , Jana Osea , Zuheng Xu , Jason Hartford

Existing cross-modal pedestrian detection (CMPD) employs complementary information from RGB and thermal-infrared (TIR) modalities to detect pedestrians in 24h-surveillance systems.RGB captures rich pedestrian details under daylight, while…

Computer Vision and Pattern Recognition · Computer Science 2026-02-09 Qian Bie , Xiao Wang , Bin Yang , Zhixi Yu , Jun Chen , Xin Xu

The Profiled Vehicle Routing Problem (PVRP) extends the classical VRP by incorporating vehicle-client-specific preferences and constraints, reflecting real-world requirements such as zone restrictions and service-level preferences. While…

Machine Learning · Computer Science 2025-08-26 Chuanbo Hua , Federico Berto , Zhikai Zhao , Jiwoo Son , Changhyun Kwon , Jinkyoo Park

Visual-inertial localization is a key problem in computer vision and robotics applications such as virtual reality, self-driving cars, and aerial vehicles. The goal is to estimate an accurate pose of an object when either the environment or…

Computer Vision and Pattern Recognition · Computer Science 2023-08-07 Felix Ott , Nisha Lakshmana Raichur , David Rügamer , Tobias Feigl , Heiko Neumann , Bernd Bischl , Christopher Mutschler

While substantial progress has been made in the absolute performance of localization and Visual Place Recognition (VPR) techniques, it is becoming increasingly clear from translating these systems into applications that other capabilities…

Computer Vision and Pattern Recognition · Computer Science 2023-07-06 Helen Carson , Jason J. Ford , Michael Milford

Multi-human parsing is an image segmentation task necessitating both instance-level and fine-grained category-level information. However, prior research has typically processed these two types of information through separate branches and…

Computer Vision and Pattern Recognition · Computer Science 2024-05-21 Jiaming Chu , Lei Jin , Junliang Xing , Jian Zhao

Multi-modal image segmentation faces real-world deployment challenges from incomplete/corrupted modalities degrading performance. While existing methods address training-inference modality gaps via specialized per-combination models, they…

Computer Vision and Pattern Recognition · Computer Science 2025-09-22 Xiaoqi Zhao , Youwei Pang , Chenyang Yu , Lihe Zhang , Huchuan Lu , Shijian Lu , Georges El Fakhri , Xiaofeng Liu

In this paper we address the task of visual place recognition (VPR), where the goal is to retrieve the correct GPS coordinates of a given query image against a huge geotagged gallery. While recent works have shown that building descriptors…

Computer Vision and Pattern Recognition · Computer Science 2022-01-26 Valerio Paolicelli , Antonio Tavera , Carlo Masone , Gabriele Berton , Barbara Caputo

LiDAR-based place recognition (LPR) is essential for global localization and loop-closure detection in large-scale SLAM systems. Existing methods typically construct global descriptors from Range Images or BEV representations for matching.…

Computer Vision and Pattern Recognition · Computer Science 2026-04-07 Shuyuan Li , Zihang Wang , Xieyuanli Chen , Wenkai Zhu , Xiaoteng Fang , Peizhou Ni , Junhao Yang , Dong Kong

Unsupervised pre-training has shown great success in skeleton-based action understanding recently. Existing works typically train separate modality-specific models, then integrate the multi-modal information for action understanding by a…

Computer Vision and Pattern Recognition · Computer Science 2023-11-07 Shengkai Sun , Daizong Liu , Jianfeng Dong , Xiaoye Qu , Junyu Gao , Xun Yang , Xun Wang , Meng Wang

Visual Place Recognition (VPR) has seen significant advances at the frontiers of matching performance and computational superiority over the past few years. However, these evaluations are performed for ground-based mobile platforms and…

Computer Vision and Pattern Recognition · Computer Science 2019-05-24 Mubariz Zaffar , Ahmad Khaliq , Shoaib Ehsan , Michael Milford , Kostas Alexis , Klaus McDonald-Maier

Recently, many multi-modal trackers prioritize RGB as the dominant modality, treating other modalities as auxiliary, and fine-tuning separately various multi-modal tasks. This imbalance in modality dependence limits the ability of methods…

Computer Vision and Pattern Recognition · Computer Science 2025-02-11 Xiantao Hu , Bineng Zhong , Qihua Liang , Zhiyi Mo , Liangtao Shi , Ying Tai , Jian Yang

Audio-visual event localization aims to localize an event that is both audible and visible in the wild, which is a widespread audio-visual scene analysis task for unconstrained videos. To address this task, we propose a Multimodal Parallel…

Computer Vision and Pattern Recognition · Computer Science 2021-04-08 Jiashuo Yu , Ying Cheng , Rui Feng
‹ Prev 1 4 5 6 7 8 10 Next ›