中文
相关论文

相关论文: MMS-VPR: Multimodal Street-Level Visual Place Reco…

200 篇论文

We present Urban-ImageNet, a large-scale multi-modal dataset and evaluation benchmark for urban space perception from user-generated social media imagery. The corpus contains over 2 Million public social media images and paired textual…

计算机视觉与模式识别 · 计算机科学 2026-05-12 Yiwei Ou , Chung Ching Cheung , Jun Yang Ang , Xiaobin Ren , Ronggui Sun , Guansong Gao , Kaiqi Zhao , Manfredo Manfredini

This paper presents an approach for creating a visual place recognition (VPR) database for localization in indoor environments from RGBD scanning sequences. The proposed approach is formulated as a minimization problem in terms of…

计算机视觉与模式识别 · 计算机科学 2024-11-01 Anastasiia Kornilova , Ivan Moskalenko , Timofei Pushkin , Fakhriddin Tojiboev , Rahim Tariverdizadeh , Gonzalo Ferrer

A cross-domain visual place recognition (VPR) task is proposed in this work, i.e., matching images of the same architectures depicted in different domains. VPR is commonly treated as an image retrieval task, where a query image from an…

计算机视觉与模式识别 · 计算机科学 2019-09-12 Ziqi Wang , Jiahui Li , Seyran Khademi , Jan van Gemert

Over the past decade, most methods in visual place recognition (VPR) have used neural networks to produce feature representations. These networks typically produce a global representation of a place image using only this image itself and…

计算机视觉与模式识别 · 计算机科学 2024-04-02 Feng Lu , Xiangyuan Lan , Lijun Zhang , Dongmei Jiang , Yaowei Wang , Chun Yuan

Visual Place Recognition (VPR) is a crucial capability for long-term autonomous robots, enabling them to identify previously visited locations using visual information. However, existing methods remain limited in indoor settings due to the…

计算机视觉与模式识别 · 计算机科学 2025-08-12 Huaqi Tao , Bingxi Liu , Calvin Chen , Tingjun Huang , He Li , Jinqiang Cui , Hong Zhang

Geo-temporal understanding, the ability to infer location, time, and contextual properties from visual input alone, underpins applications such as disaster management, traffic planning, embodied navigation, world modeling, and geography…

Depth estimation is a fundamental component of spatial perception for autonomous driving and other unmanned systems operating in open urban environments. Existing depth datasets such as KITTI, nuScenes, and DDAD have advanced the field but…

计算机视觉与模式识别 · 计算机科学 2026-05-19 Xianda Guo , Ruijun Zhang , Yiqun Duan , Ruilin Wang , Matteo Poggi , Keyuan Zhou , Wenzhao Zheng , Wenke Huang , Gangwei Xu , Yanlun Peng , Yuan Si , Qin Zou

Multi-pedestrian tracking in aerial imagery has several applications such as large-scale event monitoring, disaster management, search-and-rescue missions, and as input into predictive crowd dynamic models. Due to the challenges such as the…

计算机视觉与模式识别 · 计算机科学 2020-06-30 Maximilian Kraus , Seyed Majid Azimi , Emec Ercelik , Reza Bahmanyar , Peter Reinartz , Alois Knoll

Vision Language Place Recognition (VLVPR) enhances robot localization performance by incorporating natural language descriptions from images. By utilizing language information, VLVPR directs robot place matching, overcoming the constraint…

计算机视觉与模式识别 · 计算机科学 2025-02-24 Tianyi Shang , Zhenyu Li , Pengjie Xu , Jinwei Qiao

Visual place recognition (VPR) enables autonomous systems to localize themselves within an environment using image information. While VPR techniques built upon a Convolutional Neural Network (CNN) backbone dominate state-of-the-art VPR…

计算机视觉与模式识别 · 计算机科学 2023-12-21 Bruno Arcanjo , Bruno Ferrarini , Maria Fasli , Michael Milford , Klaus D. McDonald-Maier , Shoaib Ehsan

The rapid progress of Large Language Models (LLMs) has spurred growing interest in Multi-modal LLMs (MLLMs) and motivated the development of benchmarks to evaluate their perceptual and comprehension abilities. Existing benchmarks, however,…

计算机视觉与模式识别 · 计算机科学 2026-03-25 Purui Bai , Tao Wu , Jiayang Sun , Xinyue Liu , Huaibo Huang , Ran He

Pedestrian detection has achieved significant progress with the availability of existing benchmark datasets. However, there is a gap in the diversity and density between real world requirements and current pedestrian detection benchmarks:…

计算机视觉与模式识别 · 计算机科学 2019-09-27 Shifeng Zhang , Yiliang Xie , Jun Wan , Hansheng Xia , Stan Z. Li , Guodong Guo

Existing Moment retrieval (MR) methods focus on Single-Moment Retrieval (SMR). However, one query can correspond to multiple relevant moments in real-world applications. This makes the existing datasets and methods insufficient for video…

计算机视觉与模式识别 · 计算机科学 2025-10-21 Zhuo Cao , Heming Du , Bingqing Zhang , Xin Yu , Xue Li , Sen Wang

While various multimodal multi-image evaluation datasets have been emerged, but these datasets are primarily based on English, and there has yet to be a Chinese multi-image dataset. To fill this gap, we introduce RealBench, the first…

计算与语言 · 计算机科学 2025-09-23 Fei Zhao , Chengqiang Lu , Yufan Shen , Qimeng Wang , Yicheng Qian , Haoxin Zhang , Yan Gao , Yi Wu , Yao Hu , Zhen Wu , Shangyu Xing , Xinyu Dai

In this paper we present a large-scale visual object detection and tracking benchmark, named VisDrone2018, aiming at advancing visual understanding tasks on the drone platform. The images and video sequences in the benchmark were captured…

计算机视觉与模式识别 · 计算机科学 2018-04-24 Pengfei Zhu , Longyin Wen , Xiao Bian , Haibin Ling , Qinghua Hu

We propose the Multi-modal Untrimmed Video Retrieval task, along with a new benchmark (MUVR) to advance video retrieval for long-video platforms. MUVR aims to retrieve untrimmed videos containing relevant segments using multi-modal queries.…

计算机视觉与模式识别 · 计算机科学 2025-10-27 Yue Feng , Jinwei Hu , Qijia Lu , Jiawei Niu , Li Tan , Shuo Yuan , Ziyi Yan , Yizhen Jia , Qingzhi He , Shiping Ge , Ethan Q. Chen , Wentong Li , Limin Wang , Jie Qin

In this data article, we introduce the Multi-Modal Event-based Vehicle Detection and Tracking (MEVDT) dataset. This dataset provides a synchronized stream of event data and grayscale images of traffic scenes, captured using the Dynamic and…

计算机视觉与模式识别 · 计算机科学 2024-07-31 Zaid A. El Shair , Samir A. Rawashdeh

Visual Place Recognition (VPR) aims to retrieve frames from a geotagged database that are located at the same place as the query frame. To improve the robustness of VPR in perceptually aliasing scenarios, sequence-based VPR methods are…

计算机视觉与模式识别 · 计算机科学 2024-01-30 Junqiao Zhao , Fenglin Zhang , Yingfeng Cai , Gengxuan Tian , Wenjie Mu , Chen Ye , Tiantian Feng

There has been exciting recent progress in using radar as a sensor for robot navigation due to its increased robustness to varying environmental conditions. However, within these different radar perception systems, ground penetrating radar…

机器人学 · 计算机科学 2021-07-19 Alexander Baikovitz , Paloma Sodhi , Michael Dille , Michael Kaess

Cross-modal place recognition methods are flexible GPS-alternatives under varying environment conditions and sensor setups. However, this task is non-trivial since extracting consistent and robust global descriptors from different…

计算机视觉与模式识别 · 计算机科学 2025-03-18 Yun-Jin Li , Mariia Gladkova , Yan Xia , Rui Wang , Daniel Cremers