English
Related papers

Related papers: Human Cognitive Benchmarks Reveal Foundational Vis…

200 papers

CAPTCHA, originally designed to distinguish humans from robots, has evolved into a real-world benchmark for assessing the spatial reasoning capabilities of vision-language models. In this work, we first show that step-by-step reasoning is…

Computer Vision and Pattern Recognition · Computer Science 2025-11-18 Python Song , Luke Tenyi Chang , Yun-Yun Tsai , Penghui Li , Junfeng Yang

Large Vision-Language Models (LVLMs) suffer from hallucination issues, wherein the models generate plausible-sounding but factually incorrect outputs, undermining their reliability. A comprehensive quantitative evaluation is necessary to…

Computation and Language · Computer Science 2024-10-07 Haoyi Qiu , Wenbo Hu , Zi-Yi Dou , Nanyun Peng

Multimodal large language models (MLLMs) demonstrate strong perception and reasoning performance on existing remote sensing (RS) benchmarks. However, most prior benchmarks rely on low-resolution imagery, and some high-resolution benchmarks…

Computer Vision and Pattern Recognition · Computer Science 2025-12-22 Yunkai Dang , Meiyi Zhu , Donghao Wang , Yizhuo Zhang , Jiacheng Yang , Qi Fan , Yuekun Yang , Wenbin Li , Feng Miao , Yang Gao

We introduce Blink, a new benchmark for multimodal language models (LLMs) that focuses on core visual perception abilities not found in other evaluations. Most of the Blink tasks can be solved by humans "within a blink" (e.g., relative…

Computer Vision and Pattern Recognition · Computer Science 2024-07-04 Xingyu Fu , Yushi Hu , Bangzheng Li , Yu Feng , Haoyu Wang , Xudong Lin , Dan Roth , Noah A. Smith , Wei-Chiu Ma , Ranjay Krishna

Multimodal large language models (MLLMs) have achieved remarkable performance across diverse vision-and-language tasks. However, their potential in face recognition remains underexplored. In particular, the performance of open-source MLLMs…

Computer Vision and Pattern Recognition · Computer Science 2025-10-17 Hatef Otroshi Shahreza , Sébastien Marcel

Despite the impressive performance of vision-language models (VLMs) on downstream tasks, their ability to understand and reason about causal relationships in visual inputs remains unclear. Robust causal reasoning is fundamental to solving…

Computation and Language · Computer Science 2026-02-05 Zhaotian Weng , Haoxuan Li , Xin Eric Wang , Kuan-Hao Huang , Jieyu Zhao

Large Vision Language Models (LVLMs) excel in various vision-language tasks. Yet, their robustness to visual variations in position, scale, orientation, and context that objects in natural scenes inevitably exhibit due to changes in…

Computer Vision and Pattern Recognition · Computer Science 2025-06-03 Zhiyuan Fan , Yumeng Wang , Sandeep Polisetty , Yi R. Fung

Although Multimodal Large Language Models (MLLMs) have demonstrated increasingly impressive performance in chart understanding, most of them exhibit alarming hallucinations and significant performance degradation when handling non-annotated…

Computation and Language · Computer Science 2025-12-16 Xiao Zhang , Dongyuan Li , Liuyu Xiang , Yao Zhang , Cheng Zhong , Zhaofeng He

Current benchmarks for evaluating Vision Language Models (VLMs) often fall short in thoroughly assessing model abilities to understand and process complex visual and textual content. They typically focus on simple tasks that do not require…

Computer Vision and Pattern Recognition · Computer Science 2024-11-26 Harsha Vardhan Khurdula , Basem Rizk , Indus Khaitan , Janit Anjaria , Aviral Srivastava , Rajvardhan Khaitan

Large language models (LLMs) have shown remarkable ability in various language tasks, especially with their emergent in-context learning capability. Extending LLMs to incorporate visual inputs, large vision-language models (LVLMs) have…

Machine Learning · Computer Science 2025-10-13 Aneesh Komanduri , Karuna Bhaila , Xintao Wu

Large vision-language models (LVLMs) have demonstrated remarkable achievements, yet the generation of non-factual responses remains prevalent in fact-seeking question answering (QA). Current multimodal fact-seeking benchmarks primarily…

Computation and Language · Computer Science 2025-03-11 Yanling Wang , Yihan Zhao , Xiaodong Chen , Shasha Guo , Lixin Liu , Haoyang Li , Yong Xiao , Jing Zhang , Qi Li , Ke Xu

Significant research efforts have been made to scale and improve vision-language model (VLM) training approaches. Yet, with an ever-growing number of benchmarks, researchers are tasked with the heavy burden of implementing each protocol,…

Computer Vision and Pattern Recognition · Computer Science 2024-08-12 Haider Al-Tahan , Quentin Garrido , Randall Balestriero , Diane Bouchacourt , Caner Hazirbas , Mark Ibrahim

While Multimodal Large Language Models (MLLMs) excel on benchmarks, their processing paradigm differs from the human ability to integrate visual information. Unlike humans who naturally bridge details and high-level concepts, models tend to…

Computer Vision and Pattern Recognition · Computer Science 2025-11-25 Ming Zhong , Yuanlei Wang , Liuzhou Zhang , Arctanx An , Renrui Zhang , Hao Liang , Ming Lu , Ying Shen , Wentao Zhang

While recent Vision-Language Models (VLMs) have achieved impressive progress, it remains difficult to determine why they succeed or fail on complex reasoning tasks. Traditional benchmarks evaluate what models can answer correctly, not why…

Computer Vision and Pattern Recognition · Computer Science 2025-11-18 Ieva Bagdonaviciute , Vibhav Vineet

With the rapid advancement of Multimodal Large Language Models (MLLMs), they have demonstrated exceptional capabilities across a variety of vision-language tasks. However, current evaluation benchmarks predominantly focus on objective…

Computation and Language · Computer Science 2025-09-24 Haokun Li , Yazhou Zhang , Jizhi Ding , Qiuchi Li , Peng Zhang

We propose an efficient evaluation protocol for large vision-language models (VLMs). Given their broad knowledge and reasoning capabilities, multiple benchmarks are needed for comprehensive assessment, making evaluation computationally…

Computer Vision and Pattern Recognition · Computer Science 2025-04-15 Teppei Suzuki , Keisuke Ozawa

Vision-language models (VLMs) achieve strong benchmark results, yet can exhibit systematic perceptual weaknesses: structured, large changes to pixel values can cause confident yet nonsensical predictions, even when the underlying scene…

Computer Vision and Pattern Recognition · Computer Science 2026-03-09 Nicoleta-Nina Basoc , Adrian Cosma , Emilian Radoi

The significant advancements in visual understanding and instruction following from Multimodal Large Language Models (MLLMs) have opened up more possibilities for broader applications in diverse and universal human-centric scenarios.…

Computer Vision and Pattern Recognition · Computer Science 2024-10-10 Keliang Li , Zaifei Yang , Jiahe Zhao , Hongze Shen , Ruibing Hou , Hong Chang , Shiguang Shan , Xilin Chen

Vision-Language Models (VLMs) have demonstrated remarkable capabilities in cross-modal understanding and generation by integrating visual and textual information. While instruction tuning and parameter-efficient fine-tuning methods have…

Machine Learning · Computer Science 2025-06-12 Weiying Zheng , Ziyue Lin , Pengxin Guo , Yuyin Zhou , Feifei Wang , Liangqiong Qu

One of the primary challenges faced by deep learning is the degree to which current methods exploit superficial statistics and dataset bias, rather than learning to generalise over the specific representations they have experienced. This is…

Computer Vision and Pattern Recognition · Computer Science 2019-07-30 Damien Teney , Peng Wang , Jiewei Cao , Lingqiao Liu , Chunhua Shen , Anton van den Hengel
‹ Prev 1 8 9 10 Next ›