English
Related papers

Related papers: GeoRouter: Dynamic Paradigm Routing for Worldwide …

200 papers

Multimodal large language models (MLLMs) have exhibited remarkable performance in various visual tasks, yet still struggle with spatial reasoning. Recent efforts mitigate this by injecting geometric features from 3D foundation models, but…

Computer Vision and Pattern Recognition · Computer Science 2026-04-15 Zhaochen Liu , Limeng Qiao , Guanglu Wan , Tingting Jiang

Estimating vehicles' locations is one of the key components in intelligent traffic management systems (ITMSs) for increasing traffic scene awareness. Traditionally, stationary sensors have been employed in this regard. The development of…

Computer Vision and Pattern Recognition · Computer Science 2022-03-22 Elnaz Namazi , Rudolf Mester , Chaoru Lu , Jingyue Li

Retrieving images from the same location as a given query is an important component of multiple computer vision tasks, like Visual Place Recognition, Landmark Retrieval, Visual Localization, 3D reconstruction, and SLAM. However, existing…

Computer Vision and Pattern Recognition · Computer Science 2025-06-11 Gabriele Berton , Carlo Masone

Visual localization, i.e., camera pose estimation in a known scene, is a core component of technologies such as autonomous driving and augmented reality. State-of-the-art localization approaches often rely on image retrieval techniques for…

Computer Vision and Pattern Recognition · Computer Science 2022-06-01 Martin Humenberger , Yohann Cabon , Noé Pion , Philippe Weinzaepfel , Donghwan Lee , Nicolas Guérin , Torsten Sattler , Gabriela Csurka

Despite advances in object detection, aerial imagery remains a challenging domain, as models often fail to generalize across variations in spatial resolution, scene composition, and semantic label coverage. Differences in geographic…

Computer Vision and Pattern Recognition · Computer Science 2026-04-22 Pourya Shamsolmoali , Masoumeh Zareapoor , Michael Felsberg , Nick Pears , Yue Lu

Perceiving and reconstructing 3D scene geometry from visual inputs is crucial for autonomous driving. However, there still lacks a driving-targeted dense geometry perception model that can adapt to different scenarios and camera…

Computer Vision and Pattern Recognition · Computer Science 2025-12-19 Sicheng Zuo , Zixun Xie , Wenzhao Zheng , Shaoqing Xu , Fang Li , Shengyin Jiang , Long Chen , Zhi-Xin Yang , Jiwen Lu

Path planning for mobile robots in large dynamic environments is a challenging problem, as the robots are required to efficiently reach their given goals while simultaneously avoiding potential conflicts with other robots or dynamic…

Robotics · Computer Science 2020-09-15 Binyu Wang , Zhe Liu , Qingbiao Li , Amanda Prorok

Visual localization aims to determine the camera pose of a query image relative to a database of posed images. In recent years, deep neural networks that directly regress camera poses have gained popularity due to their fast inference…

Computer Vision and Pattern Recognition · Computer Science 2025-03-24 Siyan Dong , Shuzhe Wang , Shaohui Liu , Lulu Cai , Qingnan Fan , Juho Kannala , Yanchao Yang

Camera relocalization, a cornerstone capability of modern computer vision, accurately determines a camera's position and orientation (6-DoF) from images and is essential for applications in augmented reality (AR), mixed reality (MR),…

Computer Vision and Pattern Recognition · Computer Science 2025-07-08 Zhendong Xiao , Wu Wei , Shujie Ji , Shan Yang , Changhao Chen

Recent advancements in Large Vision-Language Models (VLMs) have shown great promise in natural image domains, allowing users to hold a dialogue about given visual content. However, such general-domain VLMs perform poorly for Remote Sensing…

Computer Vision and Pattern Recognition · Computer Science 2023-11-28 Kartik Kuckreja , Muhammad Sohail Danish , Muzammal Naseer , Abhijit Das , Salman Khan , Fahad Shahbaz Khan

The recent advancements in generative language models have demonstrated their ability to memorize knowledge from documents and recall knowledge to respond to user queries effectively. Building upon this capability, we propose to enable…

Multimedia · Computer Science 2024-02-19 Yongqi Li , Wenjie Wang , Leigang Qu , Liqiang Nie , Wenjie Li , Tat-Seng Chua

Vision-language models (VLMs) often struggle with geometric reasoning due to their limited perception of fundamental diagram elements. To tackle this challenge, we introduce GeoPerceive, a benchmark comprising diagram instances paired with…

Machine Learning · Computer Science 2026-02-27 Hao Yu , Shuning Jia , Guanghao Li , Wenhao Jiang , Chun Yuan

From a single image, visual cues can help deduce intrinsic and extrinsic camera parameters like the focal length and the gravity direction. This single-image calibration can benefit various downstream applications like image editing and 3D…

Computer Vision and Pattern Recognition · Computer Science 2024-10-18 Alexander Veicht , Paul-Edouard Sarlin , Philipp Lindenberger , Marc Pollefeys

Despite their proficiency in general tasks, Multi-modal Large Language Models (MLLMs) struggle with automatic Geometry Problem Solving (GPS), which demands understanding diagrams, interpreting symbols, and performing complex reasoning. This…

Computer Vision and Pattern Recognition · Computer Science 2025-01-13 Renqiu Xia , Mingsheng Li , Hancheng Ye , Wenjie Wu , Hongbin Zhou , Jiakang Yuan , Tianshuo Peng , Xinyu Cai , Xiangchao Yan , Bin Wang , Conghui He , Botian Shi , Tao Chen , Junchi Yan , Bo Zhang

The prevailing paradigm of perceptive humanoid locomotion relies heavily on active depth sensors. However, this depth-centric approach fundamentally discards the rich semantic and dense appearance cues of the visual world, severing…

Robotics · Computer Science 2026-03-10 Yufei Liu , Xieyuanli Chen , Hainan Pan , Chenghao Shi , Yanjie Chen , Kaihong Huang , Zhiwen Zeng , Huimin Lu

Guided image super-resolution (GISR) aims to obtain a high-resolution (HR) target image by enhancing the spatial resolution of a low-resolution (LR) target image under the guidance of a HR image. However, previous model-based methods mainly…

Image and Video Processing · Electrical Eng. & Systems 2022-03-11 Man Zhou , Keyu Yan , Jinshan Pan , Wenqi Ren , Qi Xie , Xiangyong Cao

Classifying geospatial imagery remains a major bottleneck for applications such as disaster response and land-use monitoring-particularly in regions where annotated data is scarce or unavailable. Existing tools (e.g., RS-CLIP) that claim…

Computer Vision and Pattern Recognition · Computer Science 2025-06-02 Gilles Quentin Hacheme , Girmaw Abebe Tadesse , Caleb Robinson , Akram Zaytar , Rahul Dodhia , Juan M. Lavista Ferres

Cross-View Geo-Localization (CVGL) focuses on identifying correspondences between images captured from distinct perspectives of the same geographical location. However, existing CVGL approaches are typically restricted to a single view or…

Computer Vision and Pattern Recognition · Computer Science 2026-02-03 Xudong Lu , Zhi Zheng , Yi Wan , Yongxiang Yao , Annan Wang , Renrui Zhang , Panwang Xia , Qiong Wu , Qingyun Li , Weifeng Lin , Xiangyu Zhao , Peifeng Ma , Xue Yang , Hongsheng Li

Natural-language Guided Cross-view Geo-localization (NGCG) aims to retrieve geo-tagged satellite imagery using textual descriptions of ground scenes. While recent NGCG methods commonly rely on CLIP-style dual-encoder architectures, they…

Computer Vision and Pattern Recognition · Computer Science 2026-04-14 Yuqi Chen , Xiaohan Zhang , Ahmad Arrabi , Waqas Sultani , Chen Chen , Safwan Wshah

Deep convolutional neural networks have largely benefited computer vision tasks. However, the high computational complexity limits their real-world applications. To this end, many methods have been proposed for efficient network learning,…

Computer Vision and Pattern Recognition · Computer Science 2019-12-03 Biao Qian , Yang Wang , Zhao Zhang , Richang Hong , Meng Wang , Ling Shao
‹ Prev 1 8 9 10 Next ›