中文
相关论文

相关论文: AddressVLM: Cross-view Alignment Tuning for Image …

200 篇论文

The emergence of Large Language Models (LLMs) presents unprecedented opportunities to revolutionize medical contrastive vision-language pre-training. In this paper, we show how LLMs can facilitate large-scale supervised pre-training,…

计算机视觉与模式识别 · 计算机科学 2025-09-17 Yingtai Li , Haoran Lai , Xiaoqian Zhou , Shuai Ming , Wenxin Ma , Wei Wei , Shaohua Kevin Zhou

Large vision-language models (VLMs) typically process hundreds or thousands of visual tokens per image or video frame, incurring quadratic attention cost and substantial redundancy. Existing token reduction methods often ignore the textual…

计算机视觉与模式识别 · 计算机科学 2025-12-24 Kaitong Cai , Jusheng Zhang , Jing Yang , Yijia Fan , Pengtao Xie , Jian Wang , Keze Wang

Image geolocalization, the task of identifying the geographic location depicted in an image, is important for applications in crisis response, digital forensics, and location-based intelligence. While recent advances in large language…

计算机视觉与模式识别 · 计算机科学 2026-05-19 Lingyao Li , Runlong Yu , Qikai Hu , Bowei Li , Min Deng , Yang Zhou , Xiaowei Jia

Recent document question answering models consist of two key components: the vision encoder, which captures layout and visual elements in images, and a Large Language Model (LLM) that helps contextualize questions to the image and…

计算机视觉与模式识别 · 计算机科学 2023-09-27 Nidhi Hegde , Sujoy Paul , Gagan Madan , Gaurav Aggarwal

Cross-view geo-localization (CVGL) estimates a camera's location by matching a street-view image to geo-referenced overhead imagery, enabling GPS-denied localization and navigation. Existing methods almost universally formulate CVGL as an…

计算机视觉与模式识别 · 计算机科学 2026-03-27 Yunus Talha Erzurumlu , Jiyong Kwag , Alper Yilmaz

Large Vision-Language Models (LVLMs) are pivotal for real-world AI tasks like embodied intelligence due to their strong vision-language reasoning abilities. However, current LVLMs process entire images at the token level, which is…

计算与语言 · 计算机科学 2025-05-20 Run Luo , Renke Shan , Longze Chen , Ziqiang Liu , Lu Wang , Min Yang , Xiaobo Xia

Region representation learning plays a pivotal role in urban computing by extracting meaningful features from unlabeled urban data. Analogous to how perceived facial age reflects an individual's health, the visual appearance of a city…

人工智能 · 计算机科学 2025-12-02 Yimei Zhang , Guojiang Shen , Kaili Ning , Tongwei Ren , Xuebo Qiu , Mengmeng Wang , Xiangjie Kong

Visual geo-localization demands in-depth knowledge and advanced reasoning skills to associate images with precise real-world geographic locations. Existing image database retrieval methods are limited by the impracticality of storing…

计算机视觉与模式识别 · 计算机科学 2024-10-16 Xiao Han , Chen Zhu , Xiangyu Zhao , Hengshu Zhu

Accurate visual localization in dense urban environments poses a fundamental task in photogrammetry, geospatial information science, and robotics. While imagery is a low-cost and widely accessible sensing modality, its effectiveness on…

机器人学 · 计算机科学 2025-09-10 Yandi Yang , Jianping Li , Youqi Liao , Yuhao Li , Yizhe Zhang , Zhen Dong , Bisheng Yang , Naser El-Sheimy

How well do text-only large language models (LLMs) align with the visual world? We present a systematic evaluation of this question by incorporating frozen representations of various language models into a discriminative vision-language…

计算与语言 · 计算机科学 2026-01-19 Jona Ruthardt , Gertjan J. Burghouts , Serge Belongie , Yuki M. Asano

Recent Large Vision-Language Models (LVLMs) demonstrate remarkable capabilities in image understanding and natural language generation. However, current approaches focus predominantly on global image understanding, struggling to simulate…

计算机视觉与模式识别 · 计算机科学 2026-02-25 Fan Yang , Shurong Zheng , Hongyin Zhao , Yufei Zhan , Xin Li , Yousong Zhu , Chaoyang Zhao Ming Tang , Jinqiao Wang

Vision-Language Models (VLMs) are becoming increasingly powerful, demonstrating strong performance on a variety of tasks that require both visual and textual understanding. Their strong generalisation abilities make them a promising…

计算机视觉与模式识别 · 计算机科学 2025-12-11 Nikos Theodoridis , Tim Brophy , Reenu Mohandas , Ganesh Sistu , Fiachra Collins , Anthony Scanlan , Ciaran Eising

Vision-Language Models (VLMs) have shown remarkable capabilities across diverse visual tasks, including image recognition, video understanding, and Visual Question Answering (VQA) when explicitly trained for these tasks. Despite these…

This paper proposes LayoutLLM, a more flexible document analysis method for understanding imaged documents. Visually Rich Document Understanding tasks, such as document image classification and information extraction, have gained…

计算与语言 · 计算机科学 2024-03-22 Masato Fujitake

Recent advancements in Vision-Language (VL) research have sparked new benchmarks for complex visual reasoning, challenging models' advanced reasoning ability. Traditional Vision-Language Models (VLMs) perform well in visual perception tasks…

计算机视觉与模式识别 · 计算机科学 2024-09-24 Zhiyuan Li , Dongnan Liu , Chaoyi Zhang , Heng Wang , Tengfei Xue , Weidong Cai

Generalizing an object detector trained on a single domain to multiple unseen domains is a challenging task. Existing methods typically introduce image or feature augmentation to diversify the source domain to raise the robustness of the…

计算机视觉与模式识别 · 计算机科学 2025-05-13 Hongda Qin , Xiao Lu , Zhiyong Wei , Yihong Cao , Kailun Yang , Ningjiang Chen

Vision-language models (VLMs) have demonstrated strong performance in image geolocation, a capability further sharpened by frontier multimodal large reasoning models (MLRMs). This poses a significant privacy risk, as these widely accessible…

密码学与安全 · 计算机科学 2026-02-19 Ruixin Yang , Ethan Mendes , Arthur Wang , James Hays , Sauvik Das , Wei Xu , Alan Ritter

This paper addresses the problem of cross-view image geo-localization, where the geographic location of a ground-level street-view query image is estimated by matching it against a large scale aerial map (e.g., a high-resolution satellite…

计算机视觉与模式识别 · 计算机科学 2019-11-28 Yujiao Shi , Xin Yu , Liu Liu , Tong Zhang , Hongdong Li

Cross-view geolocalization (CVGL) systems, while effective at retrieving a list of relevant candidates (high Recall@k), often fail to identify the single best match (low Top-1 accuracy). This work investigates the use of zero-shot…

计算机视觉与模式识别 · 计算机科学 2026-03-31 Yunus Talha Erzurumlu , John E. Anderson , William J. Shuart , Charles Toth , Alper Yilmaz

Building properties, such as height, usage, and material, play a crucial role in spatial data infrastructures, supporting various urban applications. Despite their importance, comprehensive building attribute data remain scarce in many…

计算机视觉与模式识别 · 计算机科学 2025-11-04 Xiucheng Liang , Jinheng Xie , Tianhong Zhao , Rudi Stouffs , Filip Biljecki