English
Related papers

Related papers: MMS-VPR: Multimodal Street-Level Visual Place Reco…

200 papers

We present Urban-ImageNet, a large-scale multi-modal dataset and evaluation benchmark for urban space perception from user-generated social media imagery. The corpus contains over 2 Million public social media images and paired textual…

Computer Vision and Pattern Recognition · Computer Science 2026-05-12 Yiwei Ou , Chung Ching Cheung , Jun Yang Ang , Xiaobin Ren , Ronggui Sun , Guansong Gao , Kaiqi Zhao , Manfredo Manfredini

This paper presents an approach for creating a visual place recognition (VPR) database for localization in indoor environments from RGBD scanning sequences. The proposed approach is formulated as a minimization problem in terms of…

Computer Vision and Pattern Recognition · Computer Science 2024-11-01 Anastasiia Kornilova , Ivan Moskalenko , Timofei Pushkin , Fakhriddin Tojiboev , Rahim Tariverdizadeh , Gonzalo Ferrer

A cross-domain visual place recognition (VPR) task is proposed in this work, i.e., matching images of the same architectures depicted in different domains. VPR is commonly treated as an image retrieval task, where a query image from an…

Computer Vision and Pattern Recognition · Computer Science 2019-09-12 Ziqi Wang , Jiahui Li , Seyran Khademi , Jan van Gemert

Over the past decade, most methods in visual place recognition (VPR) have used neural networks to produce feature representations. These networks typically produce a global representation of a place image using only this image itself and…

Computer Vision and Pattern Recognition · Computer Science 2024-04-02 Feng Lu , Xiangyuan Lan , Lijun Zhang , Dongmei Jiang , Yaowei Wang , Chun Yuan

Visual Place Recognition (VPR) is a crucial capability for long-term autonomous robots, enabling them to identify previously visited locations using visual information. However, existing methods remain limited in indoor settings due to the…

Computer Vision and Pattern Recognition · Computer Science 2025-08-12 Huaqi Tao , Bingxi Liu , Calvin Chen , Tingjun Huang , He Li , Jinqiang Cui , Hong Zhang

Geo-temporal understanding, the ability to infer location, time, and contextual properties from visual input alone, underpins applications such as disaster management, traffic planning, embodied navigation, world modeling, and geography…

Computer Vision and Pattern Recognition · Computer Science 2026-05-26 Azmine Toushik Wasi , Shahriyar Zaman Ridoy , Koushik Ahamed Tonmoy , Kinga Tshering , S. M. Muhtasimul Hasan , Wahid Faisal , Tasnim Mohiuddin , Md Rizwan Parvez

Depth estimation is a fundamental component of spatial perception for autonomous driving and other unmanned systems operating in open urban environments. Existing depth datasets such as KITTI, nuScenes, and DDAD have advanced the field but…

Computer Vision and Pattern Recognition · Computer Science 2026-05-19 Xianda Guo , Ruijun Zhang , Yiqun Duan , Ruilin Wang , Matteo Poggi , Keyuan Zhou , Wenzhao Zheng , Wenke Huang , Gangwei Xu , Yanlun Peng , Yuan Si , Qin Zou

Multi-pedestrian tracking in aerial imagery has several applications such as large-scale event monitoring, disaster management, search-and-rescue missions, and as input into predictive crowd dynamic models. Due to the challenges such as the…

Computer Vision and Pattern Recognition · Computer Science 2020-06-30 Maximilian Kraus , Seyed Majid Azimi , Emec Ercelik , Reza Bahmanyar , Peter Reinartz , Alois Knoll

Vision Language Place Recognition (VLVPR) enhances robot localization performance by incorporating natural language descriptions from images. By utilizing language information, VLVPR directs robot place matching, overcoming the constraint…

Computer Vision and Pattern Recognition · Computer Science 2025-02-24 Tianyi Shang , Zhenyu Li , Pengjie Xu , Jinwei Qiao

Visual place recognition (VPR) enables autonomous systems to localize themselves within an environment using image information. While VPR techniques built upon a Convolutional Neural Network (CNN) backbone dominate state-of-the-art VPR…

Computer Vision and Pattern Recognition · Computer Science 2023-12-21 Bruno Arcanjo , Bruno Ferrarini , Maria Fasli , Michael Milford , Klaus D. McDonald-Maier , Shoaib Ehsan

The rapid progress of Large Language Models (LLMs) has spurred growing interest in Multi-modal LLMs (MLLMs) and motivated the development of benchmarks to evaluate their perceptual and comprehension abilities. Existing benchmarks, however,…

Computer Vision and Pattern Recognition · Computer Science 2026-03-25 Purui Bai , Tao Wu , Jiayang Sun , Xinyue Liu , Huaibo Huang , Ran He

Pedestrian detection has achieved significant progress with the availability of existing benchmark datasets. However, there is a gap in the diversity and density between real world requirements and current pedestrian detection benchmarks:…

Computer Vision and Pattern Recognition · Computer Science 2019-09-27 Shifeng Zhang , Yiliang Xie , Jun Wan , Hansheng Xia , Stan Z. Li , Guodong Guo

Existing Moment retrieval (MR) methods focus on Single-Moment Retrieval (SMR). However, one query can correspond to multiple relevant moments in real-world applications. This makes the existing datasets and methods insufficient for video…

Computer Vision and Pattern Recognition · Computer Science 2025-10-21 Zhuo Cao , Heming Du , Bingqing Zhang , Xin Yu , Xue Li , Sen Wang

While various multimodal multi-image evaluation datasets have been emerged, but these datasets are primarily based on English, and there has yet to be a Chinese multi-image dataset. To fill this gap, we introduce RealBench, the first…

Computation and Language · Computer Science 2025-09-23 Fei Zhao , Chengqiang Lu , Yufan Shen , Qimeng Wang , Yicheng Qian , Haoxin Zhang , Yan Gao , Yi Wu , Yao Hu , Zhen Wu , Shangyu Xing , Xinyu Dai

In this paper we present a large-scale visual object detection and tracking benchmark, named VisDrone2018, aiming at advancing visual understanding tasks on the drone platform. The images and video sequences in the benchmark were captured…

Computer Vision and Pattern Recognition · Computer Science 2018-04-24 Pengfei Zhu , Longyin Wen , Xiao Bian , Haibin Ling , Qinghua Hu

We propose the Multi-modal Untrimmed Video Retrieval task, along with a new benchmark (MUVR) to advance video retrieval for long-video platforms. MUVR aims to retrieve untrimmed videos containing relevant segments using multi-modal queries.…

Computer Vision and Pattern Recognition · Computer Science 2025-10-27 Yue Feng , Jinwei Hu , Qijia Lu , Jiawei Niu , Li Tan , Shuo Yuan , Ziyi Yan , Yizhen Jia , Qingzhi He , Shiping Ge , Ethan Q. Chen , Wentong Li , Limin Wang , Jie Qin

In this data article, we introduce the Multi-Modal Event-based Vehicle Detection and Tracking (MEVDT) dataset. This dataset provides a synchronized stream of event data and grayscale images of traffic scenes, captured using the Dynamic and…

Computer Vision and Pattern Recognition · Computer Science 2024-07-31 Zaid A. El Shair , Samir A. Rawashdeh

Visual Place Recognition (VPR) aims to retrieve frames from a geotagged database that are located at the same place as the query frame. To improve the robustness of VPR in perceptually aliasing scenarios, sequence-based VPR methods are…

Computer Vision and Pattern Recognition · Computer Science 2024-01-30 Junqiao Zhao , Fenglin Zhang , Yingfeng Cai , Gengxuan Tian , Wenjie Mu , Chen Ye , Tiantian Feng

There has been exciting recent progress in using radar as a sensor for robot navigation due to its increased robustness to varying environmental conditions. However, within these different radar perception systems, ground penetrating radar…

Robotics · Computer Science 2021-07-19 Alexander Baikovitz , Paloma Sodhi , Michael Dille , Michael Kaess

Cross-modal place recognition methods are flexible GPS-alternatives under varying environment conditions and sensor setups. However, this task is non-trivial since extracting consistent and robust global descriptors from different…

Computer Vision and Pattern Recognition · Computer Science 2025-03-18 Yun-Jin Li , Mariia Gladkova , Yan Xia , Rui Wang , Daniel Cremers