English
Related papers

Related papers: MedSG-Bench: A Benchmark for Medical Image Sequenc…

200 papers

Recent advances in multimodal large language models (MLLMs) have led to impressive progress across various benchmarks. However, their capability in understanding infrared images remains unexplored. To address this gap, we introduce…

Computer Vision and Pattern Recognition · Computer Science 2025-12-11 Tao Zhang , Yuyang Hong , Yang Xia , Kun Ding , Zeyu Zhang , Ying Wang , Shiming Xiang , Chunhong Pan

Large language models (LLMs) have achieved strong performance on medical exam-style tasks, motivating growing interest in their deployment in real-world clinical settings. However, clinical decision-making is inherently safety-critical,…

Computation and Language · Computer Science 2026-04-13 Xiaohan Ren , Chenxiao Fan , Wenyin Ma , Hongliang He , Chongming Gao , Xiaoyan Zhao , Fuli Feng

Despite remarkable advancements in pixel-level medical image perception, existing methods are either limited to specific tasks or heavily rely on accurate bounding boxes or text labels as input prompts. However, the medical knowledge…

Computer Vision and Pattern Recognition · Computer Science 2025-04-24 Qinyue Tong , Ziqian Lu , Jun Liu , Yangming Zheng , Zheming Lu

Large Vision-Language Models (LVLMs) show significant strides in general-purpose multimodal applications such as visual dialogue and embodied navigation. However, existing multimodal evaluation benchmarks cover a limited number of…

With enhanced capabilities and widespread applications, Multimodal Large Language Models (MLLMs) are increasingly required to process and reason over multiple images simultaneously. However, existing MLLM benchmarks focus either on…

While Multimodal Large Language Models (MLLMs) excel at many vision tasks, it is unknown if they exhibit human-like perceptual behaviors. To evaluate this, we introduce HVSBench, the first large-scale benchmark with over 85,000 samples…

Computer Vision and Pattern Recognition · Computer Science 2025-12-18 Jiaying Lin , Shuquan Ye , Dan Xu , Wanli Ouyang , Rynson W. H. Lau

Comprehending text-rich visual content is paramount for the practical application of Multimodal Large Language Models (MLLMs), since text-rich scenarios are ubiquitous in the real world, which are characterized by the presence of extensive…

Computer Vision and Pattern Recognition · Computer Science 2024-04-26 Bohao Li , Yuying Ge , Yi Chen , Yixiao Ge , Ruimao Zhang , Ying Shan

Despite recent progress in text-prompt-based medical image segmentation, these methods are limited to single-round dialogues and fail to support multi-round reasoning, which is important for medical education scenarios. In this work, we…

Computer Vision and Pattern Recognition · Computer Science 2026-03-12 Qinyue Tong , Ziqian Lu , Jun Liu , Rui Zuo , Zheming Lu , Yueming Jin

Multimodal Large Language Models (MLLMs) show promising results as decision-making engines for embodied agents operating in complex, physical environments. However, existing benchmarks often prioritize high-level planning or spatial…

Computer Vision and Pattern Recognition · Computer Science 2026-03-13 Dayong Liu , Chao Xu , Weihong Chen , Suyu Zhang , Juncheng Wang , Jiankang Deng , Baigui Sun , Yang Liu

Spatial cognition is fundamental to real-world multimodal intelligence, allowing models to effectively interact with the physical environment. While multimodal large language models (MLLMs) have made significant strides, existing benchmarks…

Artificial Intelligence · Computer Science 2026-05-08 Peiran Xu , Sudong Wang , Yao Zhu , Jianing Li , Gege Qi , Yunjian Zhang

Multimodal large language models (MLLMs) have shown remarkable progress in high-level semantic tasks such as visual question answering, image captioning, and emotion recognition. However, despite advancements, there remains a lack of…

Computer Vision and Pattern Recognition · Computer Science 2025-11-17 Shezheng Song , Chengxiang He , Shan Zhao , Chengyu Wang , Qian Wan , Tianwei Yan , Meng Wang

Medical Visual Grounding (MVG) aims to identify diagnostically relevant phrases from free-text radiology reports and localize their corresponding regions in medical images, providing interpretable visual evidence to support clinical…

Computer Vision and Pattern Recognition · Computer Science 2026-04-03 Yifan Gao , Tao Zhou , Yi Zhou , Ke Zou , Yizhe Zhang , Huazhu Fu

The rapid evolution of Multi-modality Large Language Models (MLLMs) has catalyzed a shift in computer vision from specialized models to general-purpose foundation models. Nevertheless, there is still an inadequacy in assessing the abilities…

Computer Vision and Pattern Recognition · Computer Science 2024-01-02 Haoning Wu , Zicheng Zhang , Erli Zhang , Chaofeng Chen , Liang Liao , Annan Wang , Chunyi Li , Wenxiu Sun , Qiong Yan , Guangtao Zhai , Weisi Lin

Vision-Language Models (VLMs) have shown impressive performance in vision tasks, but adapting them to new domains often requires expensive fine-tuning. Prompt tuning techniques, including textual, visual, and multimodal prompting, offer…

Computer Vision and Pattern Recognition · Computer Science 2024-10-10 Rabin Adhikari , Safal Thapaliya , Manish Dhakal , Bishesh Khanal

Recent progress in Large Vision-Language Models (LVLMs) has enabled promising applications in medical tasks, such as report generation and visual question answering. However, existing benchmarks focus mainly on the final diagnostic answer,…

Computer Vision and Pattern Recognition · Computer Science 2025-10-28 Hyungyung Lee , Geon Choi , Jung-Oh Lee , Hangyul Yoon , Hyuk Gi Hong , Edward Choi

Brain imaging analysis is crucial for diagnosing and treating brain disorders, and multimodal large language models (MLLMs) are increasingly supporting it. However, current brain imaging visual question-answering (VQA) benchmarks either…

Computer Vision and Pattern Recognition · Computer Science 2025-12-29 Zhihao Peng , Cheng Wang , Shengyuan Liu , Zhiying Liang , Zanting Ye , Minjie Ju , PeterYM Woo , Yixuan Yuan

Existing vision-language understanding benchmarks largely consist of images of objects in their usual contexts. As a consequence, recent multimodal large language models can perform well with only a shallow visual understanding by relying…

Models like OpenAI-o3 pioneer visual grounded reasoning by dynamically referencing visual regions, just like human "thinking with images". However, no benchmark exists to evaluate these capabilities holistically. To bridge this gap, we…

Computer Vision and Pattern Recognition · Computer Science 2026-03-06 Haochen Wang , Xiangtai Li , Zilong Huang , Anran Wang , Jiacong Wang , Tao Zhang , Jiani Zheng , Sule Bai , Zijian Kang , Jiashi Feng , Zhuochen Wang , Zhaoxiang Zhang

Multi-modal large language models (MLLMs) have shown promise in advancing healthcare. However, most existing models remain confined to single-image understanding, which greatly limits their applicability in clinical workflows. In practice,…

Computer Vision and Pattern Recognition · Computer Science 2025-12-01 Zhen Chen , Yihang Fu , Gabriel Madera , Mauro Giuffre , Serina Applebaum , Hyunjae Kim , Hua Xu , Qingyu Chen

The ability to distinguish subtle differences between visually similar images is essential for diverse domains such as industrial anomaly detection, medical imaging, and aerial surveillance. While comparative reasoning benchmarks for…

Computer Vision and Pattern Recognition · Computer Science 2026-03-10 Minkyu Kim , Sangheon Lee , Dongmin Park
‹ Prev 1 4 5 6 7 8 10 Next ›