English
Related papers

Related papers: GEOBench-VLM: Benchmarking Vision-Language Models …

200 papers

Vision-Language Models like GPT-4, LLaVA, and CogVLM have surged in popularity recently due to their impressive performance in several vision-language tasks. Current evaluation methods, however, overlook an essential component: uncertainty,…

Computer Vision and Pattern Recognition · Computer Science 2024-02-27 Vasily Kostumov , Bulat Nutfullin , Oleg Pilipenko , Eugene Ilyushin

Vision language models (VLMs) have shown remarkable capabilities in integrating linguistic and visual reasoning but remain fundamentally limited in understanding dynamic spatiotemporal interactions. Humans effortlessly track and reason…

Computer Vision and Pattern Recognition · Computer Science 2025-08-08 Shijie Zhou , Alexander Vilesov , Xuehai He , Ziyu Wan , Shuwang Zhang , Aditya Nagachandra , Di Chang , Dongdong Chen , Xin Eric Wang , Achuta Kadambi

Vision language models (VLMs) are an exciting emerging class of language models (LMs) that have merged classic LM capabilities with those of image processing systems. However, the ways that these capabilities combine are not always…

Computation and Language · Computer Science 2024-07-03 Qiucheng Wu , Handong Zhao , Michael Saxon , Trung Bui , William Yang Wang , Yang Zhang , Shiyu Chang

Visual Question-Answering (VQA) has become key to user experience, particularly after improved generalization capabilities of Vision-Language Models (VLMs). But evaluating VLMs for an application requirement using a standardized framework…

Computer Vision and Pattern Recognition · Computer Science 2024-12-13 Neelabh Sinha , Vinija Jain , Aman Chadha

Recent advancements in large language models (LLMs) and multi-modal models (MMs) have demonstrated their remarkable capabilities in problem-solving. Yet, their proficiency in tackling geometry math problems, which necessitates an integrated…

Artificial Intelligence · Computer Science 2024-05-20 Jiaxin Zhang , Zhongzhi Li , Mingliang Zhang , Fei Yin , Chenglin Liu , Yashar Moshfeghi

Vision-Language Models (VLMs) have shown remarkable capabilities across diverse visual tasks, including image recognition, video understanding, and Visual Question Answering (VQA) when explicitly trained for these tasks. Despite these…

Computer Vision and Pattern Recognition · Computer Science 2025-03-14 Sivan Doveh , Nimrod Shabtay , Wei Lin , Eli Schwartz , Hilde Kuehne , Raja Giryes , Rogerio Feris , Leonid Karlinsky , James Glass , Assaf Arbelle , Shimon Ullman , M. Jehanzeb Mirza

In recent years, 2D Vision-Language Models (VLMs) have made significant strides in image-text understanding tasks. However, their performance in 3D spatial comprehension, which is critical for embodied intelligence, remains limited. Recent…

Computer Vision and Pattern Recognition · Computer Science 2025-03-12 Zhangyang Qi , Zhixiong Zhang , Ye Fang , Jiaqi Wang , Hengshuang Zhao

Vision-Language Models (VLMs) have recently gained attention due to their competitive performance on multiple downstream tasks, achieved by following user-input instructions. However, VLMs still exhibit several limitations in visual…

Computer Vision and Pattern Recognition · Computer Science 2025-10-23 Simone Alghisi , Gabriel Roccabruna , Massimo Rizzoli , Seyed Mahed Mousavi , Giuseppe Riccardi

Geometric understanding - including depth and height perception - is fundamental to intelligence and crucial for navigating our environment. Despite the impressive capabilities of large Vision Language Models (VLMs), it remains unclear how…

Computer Vision and Pattern Recognition · Computer Science 2025-04-28 Shehreen Azad , Yash Jain , Rishit Garg , Yogesh S Rawat , Vibhav Vineet

General-purposed embodied agents are designed to understand the users' natural instructions or intentions and act precisely to complete universal tasks. Recently, methods based on foundation models especially Vision-Language-Action models…

Multimodal Large Language Models (MLLMs) have shown remarkable proficiency on general-purpose vision-language benchmarks, reaching or even exceeding human-level performance. However, these evaluations typically rely on standard…

Computer Vision and Pattern Recognition · Computer Science 2026-02-03 Wenjin Hou , Wei Liu , Han Hu , Xiaoxiao Sun , Serena Yeung-Levy , Hehe Fan

Large Vision-Language Models (LVLMs) have become essential for advancing the integration of visual and linguistic information. However, the evaluation of LVLMs presents significant challenges as the evaluation benchmark always demands lots…

Computer Vision and Pattern Recognition · Computer Science 2025-03-07 Han Bao , Yue Huang , Yanbo Wang , Jiayi Ye , Xiangqi Wang , Xiuying Chen , Yue Zhao , Tianyi Zhou , Mohamed Elhoseiny , Xiangliang Zhang

Benchmark accuracy is often implicitly assumed to reflect grounded visual understanding in vision-language models (VLMs), yet it remains unclear to what extent such scores truly reflect reliance on visual evidence. Motivated by a surprising…

Computer Vision and Pattern Recognition · Computer Science 2026-05-25 Zixuan Lan , Luzhe Sun , Matthew R. Walter , Jiawei Zhou

Recent breakthroughs in vision-language models (VLMs) emphasize the necessity of benchmarking human preferences in real-world multimodal interactions. To address this gap, we launched WildVision-Arena (WV-Arena), an online platform that…

Computer Vision and Pattern Recognition · Computer Science 2024-06-18 Yujie Lu , Dongfu Jiang , Wenhu Chen , William Yang Wang , Yejin Choi , Bill Yuchen Lin

Humans can imagine and manipulate visual images mentally, a capability known as spatial visualization. While many multi-modal benchmarks assess reasoning on visible visual information, the ability to infer unseen relationships through…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Siting Wang , Minnan Pei , Luoyang Sun , Cheng Deng , Yuchen Li , Kun Shao , Zheng Tian , Haifeng Zhang , Jun Wang

Crop monitoring is essential for precision agriculture, but current systems lack high-level reasoning. We introduce a novel, modular framework that uses a Visual Language Model (VLM) to guide robotic task planning, interleaving input…

Robotics · Computer Science 2026-01-21 Jose Cuaran , Kendall Koe , Aditya Potnis , Naveen Kumar Uppalapati , Girish Chowdhary

Vision-Language Models (VLMs) still lack robustness in spatial intelligence, demonstrating poor performance on spatial understanding and reasoning tasks. We attribute this gap to the absence of a visual geometry learning process capable of…

Computer Vision and Pattern Recognition · Computer Science 2025-12-01 Wenbo Hu , Jingli Lin , Yilin Long , Yunlong Ran , Lihan Jiang , Yifan Wang , Chenming Zhu , Runsen Xu , Tai Wang , Jiangmiao Pang

Vision-Language Models (VLMs) face significant challenges when dealing with the diverse resolutions and aspect ratios of real-world images, as most existing models rely on fixed, low-resolution inputs. While recent studies have explored…

Computer Vision and Pattern Recognition · Computer Science 2025-06-17 Junbo Niu , Yuanhong Zheng , Ziyang Miao , Hejun Dong , Chunjiang Ge , Hao Liang , Ma Lu , Bohan Zeng , Qiahao Zheng , Conghui He , Wentao Zhang

Visual Language Models (VLMs) are now increasingly being merged with Large Language Models (LLMs) to enable new capabilities, particularly in terms of improved interactivity and open-ended responsiveness. While these are remarkable…

This paper investigates the potential of vision-language models (VLMs) to assist people with blindness and low vision (pBLV) in navigation tasks. We evaluate state-of-the-art closed-source models, including GPT-4V, GPT-4o, Gemini-1.5-Pro,…

Computer Vision and Pattern Recognition · Computer Science 2026-03-18 Yu Li , Yuchen Zheng , Giles Hamilton-Fletcher , Marco Mezzavilla , Yao Wang , Sundeep Rangan , Maurizio Porfiri , Zhou Yu , John-Ross Rizzo
‹ Prev 1 4 5 6 7 8 10 Next ›