English
Related papers

Related papers: AddressVLM: Cross-view Alignment Tuning for Image …

200 papers

The emergence of Large Language Models (LLMs) presents unprecedented opportunities to revolutionize medical contrastive vision-language pre-training. In this paper, we show how LLMs can facilitate large-scale supervised pre-training,…

Computer Vision and Pattern Recognition · Computer Science 2025-09-17 Yingtai Li , Haoran Lai , Xiaoqian Zhou , Shuai Ming , Wenxin Ma , Wei Wei , Shaohua Kevin Zhou

Large vision-language models (VLMs) typically process hundreds or thousands of visual tokens per image or video frame, incurring quadratic attention cost and substantial redundancy. Existing token reduction methods often ignore the textual…

Computer Vision and Pattern Recognition · Computer Science 2025-12-24 Kaitong Cai , Jusheng Zhang , Jing Yang , Yijia Fan , Pengtao Xie , Jian Wang , Keze Wang

Image geolocalization, the task of identifying the geographic location depicted in an image, is important for applications in crisis response, digital forensics, and location-based intelligence. While recent advances in large language…

Computer Vision and Pattern Recognition · Computer Science 2026-05-19 Lingyao Li , Runlong Yu , Qikai Hu , Bowei Li , Min Deng , Yang Zhou , Xiaowei Jia

Recent document question answering models consist of two key components: the vision encoder, which captures layout and visual elements in images, and a Large Language Model (LLM) that helps contextualize questions to the image and…

Computer Vision and Pattern Recognition · Computer Science 2023-09-27 Nidhi Hegde , Sujoy Paul , Gagan Madan , Gaurav Aggarwal

Cross-view geo-localization (CVGL) estimates a camera's location by matching a street-view image to geo-referenced overhead imagery, enabling GPS-denied localization and navigation. Existing methods almost universally formulate CVGL as an…

Computer Vision and Pattern Recognition · Computer Science 2026-03-27 Yunus Talha Erzurumlu , Jiyong Kwag , Alper Yilmaz

Large Vision-Language Models (LVLMs) are pivotal for real-world AI tasks like embodied intelligence due to their strong vision-language reasoning abilities. However, current LVLMs process entire images at the token level, which is…

Computation and Language · Computer Science 2025-05-20 Run Luo , Renke Shan , Longze Chen , Ziqiang Liu , Lu Wang , Min Yang , Xiaobo Xia

Region representation learning plays a pivotal role in urban computing by extracting meaningful features from unlabeled urban data. Analogous to how perceived facial age reflects an individual's health, the visual appearance of a city…

Artificial Intelligence · Computer Science 2025-12-02 Yimei Zhang , Guojiang Shen , Kaili Ning , Tongwei Ren , Xuebo Qiu , Mengmeng Wang , Xiangjie Kong

Visual geo-localization demands in-depth knowledge and advanced reasoning skills to associate images with precise real-world geographic locations. Existing image database retrieval methods are limited by the impracticality of storing…

Computer Vision and Pattern Recognition · Computer Science 2024-10-16 Xiao Han , Chen Zhu , Xiangyu Zhao , Hengshu Zhu

Accurate visual localization in dense urban environments poses a fundamental task in photogrammetry, geospatial information science, and robotics. While imagery is a low-cost and widely accessible sensing modality, its effectiveness on…

Robotics · Computer Science 2025-09-10 Yandi Yang , Jianping Li , Youqi Liao , Yuhao Li , Yizhe Zhang , Zhen Dong , Bisheng Yang , Naser El-Sheimy

How well do text-only large language models (LLMs) align with the visual world? We present a systematic evaluation of this question by incorporating frozen representations of various language models into a discriminative vision-language…

Computation and Language · Computer Science 2026-01-19 Jona Ruthardt , Gertjan J. Burghouts , Serge Belongie , Yuki M. Asano

Recent Large Vision-Language Models (LVLMs) demonstrate remarkable capabilities in image understanding and natural language generation. However, current approaches focus predominantly on global image understanding, struggling to simulate…

Computer Vision and Pattern Recognition · Computer Science 2026-02-25 Fan Yang , Shurong Zheng , Hongyin Zhao , Yufei Zhan , Xin Li , Yousong Zhu , Chaoyang Zhao Ming Tang , Jinqiao Wang

Vision-Language Models (VLMs) are becoming increasingly powerful, demonstrating strong performance on a variety of tasks that require both visual and textual understanding. Their strong generalisation abilities make them a promising…

Computer Vision and Pattern Recognition · Computer Science 2025-12-11 Nikos Theodoridis , Tim Brophy , Reenu Mohandas , Ganesh Sistu , Fiachra Collins , Anthony Scanlan , Ciaran Eising

Vision-Language Models (VLMs) have shown remarkable capabilities across diverse visual tasks, including image recognition, video understanding, and Visual Question Answering (VQA) when explicitly trained for these tasks. Despite these…

Computer Vision and Pattern Recognition · Computer Science 2025-03-14 Sivan Doveh , Nimrod Shabtay , Wei Lin , Eli Schwartz , Hilde Kuehne , Raja Giryes , Rogerio Feris , Leonid Karlinsky , James Glass , Assaf Arbelle , Shimon Ullman , M. Jehanzeb Mirza

This paper proposes LayoutLLM, a more flexible document analysis method for understanding imaged documents. Visually Rich Document Understanding tasks, such as document image classification and information extraction, have gained…

Computation and Language · Computer Science 2024-03-22 Masato Fujitake

Recent advancements in Vision-Language (VL) research have sparked new benchmarks for complex visual reasoning, challenging models' advanced reasoning ability. Traditional Vision-Language Models (VLMs) perform well in visual perception tasks…

Computer Vision and Pattern Recognition · Computer Science 2024-09-24 Zhiyuan Li , Dongnan Liu , Chaoyi Zhang , Heng Wang , Tengfei Xue , Weidong Cai

Generalizing an object detector trained on a single domain to multiple unseen domains is a challenging task. Existing methods typically introduce image or feature augmentation to diversify the source domain to raise the robustness of the…

Computer Vision and Pattern Recognition · Computer Science 2025-05-13 Hongda Qin , Xiao Lu , Zhiyong Wei , Yihong Cao , Kailun Yang , Ningjiang Chen

Vision-language models (VLMs) have demonstrated strong performance in image geolocation, a capability further sharpened by frontier multimodal large reasoning models (MLRMs). This poses a significant privacy risk, as these widely accessible…

Cryptography and Security · Computer Science 2026-02-19 Ruixin Yang , Ethan Mendes , Arthur Wang , James Hays , Sauvik Das , Wei Xu , Alan Ritter

This paper addresses the problem of cross-view image geo-localization, where the geographic location of a ground-level street-view query image is estimated by matching it against a large scale aerial map (e.g., a high-resolution satellite…

Computer Vision and Pattern Recognition · Computer Science 2019-11-28 Yujiao Shi , Xin Yu , Liu Liu , Tong Zhang , Hongdong Li

Cross-view geolocalization (CVGL) systems, while effective at retrieving a list of relevant candidates (high Recall@k), often fail to identify the single best match (low Top-1 accuracy). This work investigates the use of zero-shot…

Computer Vision and Pattern Recognition · Computer Science 2026-03-31 Yunus Talha Erzurumlu , John E. Anderson , William J. Shuart , Charles Toth , Alper Yilmaz

Building properties, such as height, usage, and material, play a crucial role in spatial data infrastructures, supporting various urban applications. Despite their importance, comprehensive building attribute data remain scarce in many…

Computer Vision and Pattern Recognition · Computer Science 2025-11-04 Xiucheng Liang , Jinheng Xie , Tianhong Zhao , Rudi Stouffs , Filip Biljecki