English
Related papers

Related papers: OddGridBench: Exposing the Lack of Fine-Grained Vi…

200 papers

The ability to distinguish subtle differences between visually similar images is essential for diverse domains such as industrial anomaly detection, medical imaging, and aerial surveillance. While comparative reasoning benchmarks for…

Computer Vision and Pattern Recognition · Computer Science 2026-03-10 Minkyu Kim , Sangheon Lee , Dongmin Park

Multimodal large language models (MLLMs) hold promise for integrating diverse data modalities, but current medical adaptations such as LLaVA-Med often fail to fully exploit the synergy between color fundus photography (CFP) and optical…

Recent advancements in Vision-Language (VL) models have sparked interest in their deployment on edge devices, yet challenges in handling diverse visual modalities, manual annotation, and computational constraints remain. We introduce…

Computer Vision and Pattern Recognition · Computer Science 2024-10-02 Kaiwen Cai , Zhekai Duan , Gaowen Liu , Charles Fleming , Chris Xiaoxuan Lu

The rapid advancement of native multi-modal models and omni-models, exemplified by GPT-4o, Gemini, and o3, with their capability to process and generate content across modalities such as text and images, marks a significant milestone in the…

Computer Vision and Pattern Recognition · Computer Science 2025-05-26 Meng-Hao Guo , Xuanyu Chu , Qianrui Yang , Zhe-Han Mo , Yiqing Shen , Pei-lin Li , Xinjie Lin , Jinnian Zhang , Xin-Sheng Chen , Yi Zhang , Kiyohiro Nakayama , Zhengyang Geng , Houwen Peng , Han Hu , Shi-Min Hu

Recently, large multimodal models have built a bridge from visual to textual information, but they tend to underperform in remote sensing scenarios. This underperformance is due to the complex distribution of objects and the significant…

Computer Vision and Pattern Recognition · Computer Science 2024-06-10 Cong Yang , Zuchao Li , Lefei Zhang

Evaluating the robustness of Large Vision-Language Models (LVLMs) is essential for their continued development and responsible deployment in real-world applications. However, existing robustness benchmarks typically focus on hallucination…

Computer Vision and Pattern Recognition · Computer Science 2026-05-28 Huiyi Chen , Jiawei Peng , Dehai Min , Changchang Sun , Kaijie Chen , Yan Yan , Xu Yang , Lu Cheng

As vision-language models (VLMs) are deployed globally, their ability to understand culturally situated knowledge becomes essential. Yet, existing evaluations largely assess static recall or isolated visual grounding, leaving unanswered…

Computer Vision and Pattern Recognition · Computer Science 2026-01-27 Bryan Chen Zhengyu Tan , Zheng Weihua , Zhengyuan Liu , Nancy F. Chen , Hwaran Lee , Kenny Tsu Wei Choo , Roy Ka-Wei Lee

Vision-Language Models (VLMs) building upon the foundation of powerful large language models have made rapid progress in reasoning across visual and textual data. While VLMs perform well on vision tasks that they are trained on, our results…

Computer Vision and Pattern Recognition · Computer Science 2025-06-26 Zixuan Wu , Yoolim Kim , Carolyn Jane Anderson

Can warping tokens, rather than pixels, help multimodal large language models (MLLMs) understand how a scene appears from a nearby viewpoint? While MLLMs perform well on visual reasoning, they remain fragile to viewpoint changes, as…

Computer Vision and Pattern Recognition · Computer Science 2026-04-06 Phillip Y. Lee , Chanho Park , Mingue Park , Seungwoo Yoo , Juil Koo , Minhyuk Sung

Recent advances in multimodal large language models (MLLMs) have yielded increasingly powerful models, yet their perceptual capacities remain poorly characterized. In practice, most model families scale language component while reusing…

Computer Vision and Pattern Recognition · Computer Science 2025-12-19 Tejas Anvekar , Fenil Bardoliya , Pavan K. Turaga , Chitta Baral , Vivek Gupta

Large Vision-Language Models (LVLMs) have achieved remarkable proficiency in explicit visual recognition, effectively describing what is directly visible in an image. However, a critical cognitive gap emerges when the visual input serves…

Computer Vision and Pattern Recognition · Computer Science 2026-04-03 Seyed Amir Kasaei , Arash Marioriyad , Mahbod Khaleti , MohammadAmin Fazli , Mahdieh Soleymani Baghshah , Mohammad Hossein Rohban

State-of-the-art large multi-modal models (LMMs) face challenges when processing high-resolution images, as these inputs are converted into enormous visual tokens, many of which are irrelevant to the downstream task. In this paper, we…

Computer Vision and Pattern Recognition · Computer Science 2026-04-21 Xinyu Huang , Yuhao Dong , Weiwei Tian , Bo Li , Rui Feng , Ziwei Liu

Multimodal Large Language Models (MLLMs) have made significant advancements, demonstrating powerful capabilities in processing and understanding multimodal data. Fine-tuning MLLMs with Federated Learning (FL) allows for expanding the…

Machine Learning · Computer Science 2025-03-11 Binqian Xu , Xiangbo Shu , Haiyang Mei , Guosen Xie , Basura Fernando , Jinhui Tang

Understanding multi-image, multi-turn scenarios is a critical yet underexplored capability for Large Vision-Language Models (LVLMs). Existing benchmarks predominantly focus on static or horizontal comparisons -- e.g., spotting visual…

Computer Vision and Pattern Recognition · Computer Science 2025-12-09 Wenbo Lyu , Yingjun Du , Jinglin Zhao , Xianton Zhen , Ling Shao

The popularity of multimodal large language models (MLLMs) has triggered a recent surge in research efforts dedicated to evaluating these models. Nevertheless, existing evaluation studies of MLLMs primarily focus on the comprehension and…

Computation and Language · Computer Science 2023-10-16 Xiaocui Yang , Wenfang Wu , Shi Feng , Ming Wang , Daling Wang , Yang Li , Qi Sun , Yifei Zhang , Xiaoming Fu , Soujanya Poria

This paper proposes an introspective deep metric learning (IDML) framework for uncertainty-aware comparisons of images. Conventional deep metric learning methods focus on learning a discriminative embedding to describe the semantic features…

Computer Vision and Pattern Recognition · Computer Science 2023-09-20 Chengkun Wang , Wenzhao Zheng , Zheng Zhu , Jie Zhou , Jiwen Lu

We investigated visual reasoning limitations of both multimodal large language models (MLLMs) and image generation models (IGMs) by creating a novel benchmark to systematically compare failure modes across image-to-text and text-to-image…

Computer Vision and Pattern Recognition · Computer Science 2026-01-27 Aahana Basappa , Pranay Goel , Anusri Karra , Anish Karra , Asa Gilmore , Kevin Zhu

Large Language Models (LLMs) often exhibit implicit biases and discriminatory tendencies that reflect underlying social stereotypes. While recent alignment techniques such as RLHF and DPO have mitigated some of these issues, they remain…

Computation and Language · Computer Science 2025-11-11 Deng Yixuan , Ji Xiaoqiang

Missing-modality information on e-commerce platforms, such as absent product images or textual descriptions, often arises from annotation errors or incomplete metadata, impairing both product presentation and downstream applications such as…

Multimedia · Computer Science 2026-01-29 Junchen Fu , Wenhao Deng , Kaiwen Zheng , Ioannis Arapakis , Yu Ye , Yongxin Ni , Joemon M. Jose , Xuri Ge

Multimodal Large Language Models (MLLMs) have achieved significant advances in integrating visual and linguistic information, yet their ability to reason about complex and real-world scenarios remains limited. The existing benchmarks are…

‹ Prev 1 8 9 10 Next ›