English
Related papers

Related papers: Are vision-language models ready to zero-shot repl…

200 papers

Several recent works seek to develop foundation models specifically for medical applications, adapting general-purpose large language models (LLMs) and vision-language models (VLMs) via continued pretraining on publicly available biomedical…

Computation and Language · Computer Science 2024-11-21 Daniel P. Jeong , Saurabh Garg , Zachary C. Lipton , Michael Oberst

Recently, vision-language pretraining has emerged as a transformative technique that integrates the strengths of both visual and textual modalities, resulting in powerful vision-language models (VLMs). Leveraging web-scale pretraining data,…

Computer Vision and Pattern Recognition · Computer Science 2025-07-01 Xinyao Li , Jingjing Li , Fengling Li , Lei Zhu , Yang Yang , Heng Tao Shen

The reliable analysis of blood reports is important for health knowledge, but individuals often struggle with interpretation, leading to anxiety and overlooked issues. We explore the potential of general-purpose Vision-Language Models…

Computer Vision and Pattern Recognition · Computer Science 2026-05-01 Nadia Bakhsheshi , Hamid Beigy

This study evaluates the capabilities of Multimodal Large Language Models (LLMs) and Vision Language Models (VLMs) in the task of single-label classification of Christian Iconography. The goal was to assess whether general-purpose VLMs…

Computer Vision and Pattern Recognition · Computer Science 2025-09-24 Gianmarco Spinaci , Lukas Klic , Giovanni Colavizza

Most visual recognition studies rely heavily on crowd-labelled data in deep neural networks (DNNs) training, and they usually train a DNN for each single visual recognition task, leading to a laborious and time-consuming visual recognition…

Computer Vision and Pattern Recognition · Computer Science 2024-02-19 Jingyi Zhang , Jiaxing Huang , Sheng Jin , Shijian Lu

In regions of the Middle East and North Africa (MENA), there is a high demand for wastewater treatment plants (WWTPs), crucial for sustainable water management. Precise identification of WWTPs from satellite images enables environmental…

Computer Vision and Pattern Recognition · Computer Science 2025-12-17 Akila Premarathna , Kanishka Hewageegana , Garcia Andarcia Mariangel

We introduce VLM-Lens, a toolkit designed to enable systematic benchmarking, analysis, and interpretation of vision-language models (VLMs) by supporting the extraction of intermediate outputs from any layer during the forward pass of…

Computation and Language · Computer Science 2025-10-03 Hala Sheta , Eric Huang , Shuyu Wu , Ilia Alenabi , Jiajun Hong , Ryker Lin , Ruoxi Ning , Daniel Wei , Jialin Yang , Jiawei Zhou , Ziqiao Ma , Freda Shi

Vision-Language Models (VLMs), such as recent Qwen and Gemini models, are positioned as general-purpose AI systems capable of reasoning across domains. Yet their capabilities in scientific imaging, especially on unfamiliar and potentially…

Instrumentation and Methods for Astrophysics · Physics 2025-11-13 Mariia Drozdova , Erica Lastufka , Vitaliy Kinakh , Taras Holotyak , Daniel Schaerer , Slava Voloshynovskiy

Several recent works seek to adapt general-purpose large language models (LLMs) and vision-language models (VLMs) for medical applications through continued pretraining on publicly available biomedical corpora. These works typically claim…

Computation and Language · Computer Science 2025-07-01 Daniel P. Jeong , Pranav Mani , Saurabh Garg , Zachary C. Lipton , Michael Oberst

We demonstrate that vision language models (VLMs) are capable of recognizing the content in audio recordings when given corresponding spectrogram images. Specifically, we instruct VLMs to perform audio classification tasks in a few-shot…

Sound · Computer Science 2024-11-20 Satvik Dixit , Laurie M. Heller , Chris Donahue

Recently, large language models (LLMs) have taken the spotlight in natural language processing. Further, integrating LLMs with vision enables the users to explore emergent abilities with multimodal data. Visual language models (VLMs), such…

Computer Vision and Pattern Recognition · Computer Science 2024-02-23 Minh-Hao Van , Prateek Verma , Xintao Wu

Vision Language Models (VLMs) demonstrate promising chart comprehension capabilities. Yet, prior explorations of their visualization literacy have been limited to assessing their response correctness and fail to explore their internal…

Human-Computer Interaction · Computer Science 2025-04-09 Lianghan Dong , Anamaria Crisan

Agricultural disease management in developing countries such as India, Kenya, and Nigeria faces significant challenges due to limited access to expert plant pathologists, unreliable internet connectivity, and cost constraints that hinder…

Computer Vision and Pattern Recognition · Computer Science 2025-10-28 Mihir Gupta , Pratik Desai , Ross Greer

Language provides a natural interface to specify and evaluate performance on visual tasks. To realize this possibility, vision language models (VLMs) must successfully integrate visual and linguistic information. Our work compares VLMs to a…

Computer Vision and Pattern Recognition · Computer Science 2025-06-10 Stephanie Fu , Tyler Bonnen , Devin Guillory , Trevor Darrell

In the rapidly evolving field of artificial intelligence (AI), the application of large language models (LLMs) in agriculture, particularly in pest management, remains nascent. We aimed to prove the feasibility by evaluating the content of…

Computation and Language · Computer Science 2024-03-19 Shanglong Yang , Zhipeng Yuan , Shunbao Li , Ruoling Peng , Kang Liu , Po Yang

Effective cross-modal retrieval is essential for applications like information retrieval and recommendation systems, particularly in specialized domains such as manufacturing, where product information often consists of visual samples…

Computer Vision and Pattern Recognition · Computer Science 2025-07-08 Francesco Giuliari , Asif Khan Pattan , Mohamed Lamine Mekhalfi , Fabio Poiesi

The escalating intensity and frequency of wildfires demand innovative computational methods for rapid and accurate property damage assessment. Traditional methods are often time-consuming, while modern computer vision approaches typically…

Computer Vision and Pattern Recognition · Computer Science 2026-04-07 Miguel Esparza , Archit Gupta , Kai Yin , Yiming Xiao , Ali Mostafavi

Vision Language Models (VLMs) are pivotal for advancing perception in intelligent agents. Yet, evaluation of VLMs remains limited to predominantly English-centric benchmarks in which the image-text pairs comprise short texts. To evaluate…

Computation and Language · Computer Science 2025-10-16 Jesse Atuhurra , Iqra Ali , Tomoya Iwakura , Hidetaka Kamigaito , Tatsuya Hiraoka

Automatic dietary assessment based on food images remains a challenge, requiring precise food detection, segmentation, and classification. Vision-Language Models (VLMs) offer new possibilities by integrating visual and textual reasoning. In…