English
Related papers

Related papers: MedVision: Dataset and Benchmark for Quantitative …

200 papers

Real-world clinical practice demands multi-image comparative reasoning, yet current medical benchmarks remain limited to single-frame interpretation. We present MedFrameQA, the first benchmark explicitly designed to test multi-image medical…

Computer Vision and Pattern Recognition · Computer Science 2026-02-04 Suhao Yu , Haojin Wang , Juncheng Wu , Luyang Luo , Jingshen Wang , Cihang Xie , Pranav Rajpurkar , Carl Yang , Yang Yang , Kang Wang , Yannan Yu , Yuyin Zhou

Evaluating large language models (LLMs) in medicine is crucial because medical applications require high accuracy with little room for error. Current medical benchmarks have three main types: medical exam-based, comprehensive medical, and…

Is basic visual understanding really solved in state-of-the-art VLMs? We present VisualOverload, a slightly different visual question answering (VQA) benchmark comprising 2,720 question-answer pairs, with privately held ground-truth…

Computer Vision and Pattern Recognition · Computer Science 2026-05-26 Paul Gavrikov , Wei Lin , M. Jehanzeb Mirza , Soumya Jahagirdar , Muhammad Huzaifa , Sivan Doveh , Serena Yeung-Levy , James Glass , Hilde Kuehne

Existing medical reasoning benchmarks for vision-language models primarily focus on analyzing a patient's condition based on an image from a single visit. However, this setting deviates significantly from real-world clinical practice, where…

Computer Vision and Pattern Recognition · Computer Science 2025-09-30 Junyi Zhang , Jia-Chen Gu , Wenbo Hu , Yu Zhou , Robinson Piramuthu , Nanyun Peng

MLLMs MLLMs are beginning to appear in clinical workflows, but their ability to perform complex medical reasoning remains unclear. We present Med-CMR, a fine-grained Medical Complex Multimodal Reasoning benchmark. Med-CMR distinguishes from…

Accurate disease interpretation from radiology remains challenging due to imaging heterogeneity. Achieving expert-level diagnostic decisions requires integration of subtle image features with clinical knowledge. Yet major vision-language…

Computer Vision and Pattern Recognition · Computer Science 2025-11-12 Difei Gu , Yunhe Gao , Mu Zhou , Dimitris Metaxas

In recent years, large language models (LLMs) have demonstrated remarkable potential across various medical applications. Building on this foundation, multimodal large language models (MLLMs) integrate LLMs with visual models to process…

Computation and Language · Computer Science 2025-03-11 Xiaoyi Liang , Mouxiao Bian , Moxin Chen , Lihao Liu , Junjun He , Jie Xu , Lin Li

Endoscopic procedures are essential for diagnosing and treating internal diseases, and multi-modal large language models (MLLMs) are increasingly applied to assist in endoscopy analysis. However, current benchmarks are limited, as they…

Computer Vision and Pattern Recognition · Computer Science 2025-09-25 Shengyuan Liu , Boyun Zheng , Wenting Chen , Zhihao Peng , Zhenfei Yin , Jing Shao , Jiancong Hu , Yixuan Yuan

Multimodal Large Language Models (MLLMs) have demonstrated remarkable effectiveness in various general-domain scenarios, such as visual question answering and image captioning. Recently, researchers have increasingly focused on empowering…

Computer Vision and Pattern Recognition · Computer Science 2025-11-04 Yan Shu , Chi Liu , Robin Chen , Derek Li , Bryan Dai

Significant research efforts have been made to scale and improve vision-language model (VLM) training approaches. Yet, with an ever-growing number of benchmarks, researchers are tasked with the heavy burden of implementing each protocol,…

Computer Vision and Pattern Recognition · Computer Science 2024-08-12 Haider Al-Tahan , Quentin Garrido , Randall Balestriero , Diane Bouchacourt , Caner Hazirbas , Mark Ibrahim

The ability to distinguish subtle differences between visually similar images is essential for diverse domains such as industrial anomaly detection, medical imaging, and aerial surveillance. While comparative reasoning benchmarks for…

Computer Vision and Pattern Recognition · Computer Science 2026-03-10 Minkyu Kim , Sangheon Lee , Dongmin Park

Vision language models (VLMs) achieve strong performance on general image understanding but struggle to think with medical images, especially when performing multi-step reasoning through iterative visual interaction. Medical VLMs often rely…

Computer Vision and Pattern Recognition · Computer Science 2026-01-13 Meng Lu , Yuxing Lu , Yuchen Zhuang , Megan Mullins , Yang Xie , Guanghua Xiao , Charles Fleming , Wenqi Shi , Xuan Wang

Large language models (LLMs) have demonstrated significant capabilities in mathematical reasoning, particularly with text-based mathematical problems. However, current multi-modal large language models (MLLMs), especially those specialized…

Computation and Language · Computer Science 2024-12-03 Zhen Yang , Jinhao Chen , Zhengxiao Du , Wenmeng Yu , Weihan Wang , Wenyi Hong , Zhihuan Jiang , Bin Xu , Jie Tang

As Vision-Language Models (VLMs) advance, human-centered Assistive Technologies (ATs) for helping People with Visual Impairments (PVIs) are evolving into generalists, capable of performing multiple tasks simultaneously. However,…

Computer Vision and Pattern Recognition · Computer Science 2024-11-26 Xin Jiang , Junwei Zheng , Ruiping Liu , Jiahang Li , Jiaming Zhang , Sven Matthiesen , Rainer Stiefelhagen

Translating natural language to visualization (NL2VIS) has shown great promise for visual data analysis, but it remains a challenging task that requires multiple low-level implementations, such as natural language processing and…

Human-Computer Interaction · Computer Science 2024-08-08 Nan Chen , Yuge Zhang , Jiahang Xu , Kan Ren , Yuqing Yang

Vision-language models (VLMs) are increasingly important in medical applications; however, their evaluation in dermatology remains limited by datasets that focus primarily on image-level classification tasks such as lesion recognition.…

Computer Vision and Pattern Recognition · Computer Science 2026-01-21 Abdurrahim Yilmaz , Ozan Erdem , Ece Gokyayla , Ayda Acar , Burc Bugra Dagtas , Dilara Ilhan Erdil , Gulsum Gencoglan , Burak Temelkuran

Localized image captioning has made significant progress with models like the Describe Anything Model (DAM), which can generate detailed region-specific descriptions without explicit region-text supervision. However, such capabilities have…

Computer Vision and Pattern Recognition · Computer Science 2025-05-27 Xi Xiao , Yunbei Zhang , Thanh-Huy Nguyen , Ba-Thinh Lam , Janet Wang , Lin Zhao , Jihun Hamm , Tianyang Wang , Xingjian Li , Xiao Wang , Hao Xu , Tianming Liu , Min Xu

Medical Visual Question Answering (MedVQA) presents a significant opportunity to enhance diagnostic accuracy and healthcare delivery by leveraging artificial intelligence to interpret and answer questions based on medical images. In this…

Computer Vision and Pattern Recognition · Computer Science 2024-09-10 Xiaoman Zhang , Chaoyi Wu , Ziheng Zhao , Weixiong Lin , Ya Zhang , Yanfeng Wang , Weidi Xie

Vision-Language Models (VLMs) have shown significant potential in surgical scene analysis, yet existing models are limited by frame-level datasets and lack high-quality video data with procedural surgical knowledge. To address these…

Other Quantitative Biology · Quantitative Biology 2026-01-21 Yaoqian Li , Xikai Yang , Dunyuan Xu , Yang Yu , Litao Zhao , Xiaowei Hu , Jinpeng Li , Pheng-Ann Heng

Vision-and-language models (VLMs) have been increasingly explored in the medical domain, particularly following the success of CLIP in general domain. However, unlike the relatively straightforward pairing of 2D images and text, curating…

Computer Vision and Pattern Recognition · Computer Science 2025-08-19 Ziyang Zhang , Yang Yu , Xulei Yang , Si Yong Yeo