English
Related papers

Related papers: TemMed-Bench: Evaluating Temporal Medical Image Re…

200 papers

Recent research suggests that Vision Language Models (VLMs) often rely on inherent biases learned during training when responding to queries about visual properties of images. These biases are exacerbated when VLMs are asked highly specific…

Computer Vision and Pattern Recognition · Computer Science 2026-04-06 Saurav Sengupta , Nazanin Moradinasab , Jiebei Liu , Donald E. Brown

Recent extensive works have demonstrated that by introducing long CoT, the capabilities of MLLMs to solve complex problems can be effectively enhanced. However, the reasons for the effectiveness of such paradigms remain unclear. It is…

Computer Vision and Pattern Recognition · Computer Science 2025-07-21 Jiachen Yu , Yufei Zhan , Ziheng Wu , Yousong Zhu , Jinqiao Wang , Minghui Qiu

Vision-threatening eye diseases pose a major global health burden, with timely diagnosis limited by workforce shortages and restricted access to specialized care. While multimodal large language models (MLLMs) show promise for medical image…

The rapid evolution of video generative models has shifted their focus from producing visually plausible outputs to tackling tasks requiring physical plausibility and logical consistency. However, despite recent breakthroughs such as Veo…

Computer Vision and Pattern Recognition · Computer Science 2025-12-23 Harold Haodong Chen , Disen Lan , Wen-Jie Shu , Qingyang Liu , Zihan Wang , Sirui Chen , Wenkai Cheng , Kanghao Chen , Hongfei Zhang , Zixin Zhang , Rongjin Guo , Yu Cheng , Ying-Cong Chen

General-purpose large Vision-Language Models (VLMs) demonstrate strong capabilities in generating detailed descriptions for natural images. However, their performance in the medical domain remains suboptimal, even for relatively…

Computer Vision and Pattern Recognition · Computer Science 2025-10-23 Yifan Li , Fenghe Tang , Yingtai Li , Shaohua Kevin Zhou

Spatio-temporal reasoning is a core capability for Multimodal Large Language Models (MLLMs) operating in the real world. As such, evaluating it precisely has become an essential challenge. However, existing spatio-temporal reasoning…

Computer Vision and Pattern Recognition · Computer Science 2026-05-22 Jinho Park , Youbin Kim , Hogun Park , Eunbyung Park

The recent success of large language and vision models (LLVMs) on vision question answering (VQA), particularly their applications in medicine (Med-VQA), has shown a great potential of realizing effective visual assistants for healthcare.…

Computation and Language · Computer Science 2024-04-04 Jinge Wu , Yunsoo Kim , Honghan Wu

Large language models (LLMs) have demonstrated immense capabilities in understanding textual data and are increasingly being adopted to help researchers accelerate scientific discovery through knowledge extraction (information retrieval),…

Computer Vision and Pattern Recognition · Computer Science 2025-05-30 Robinson Umeike , Neil Getty , Fangfang Xia , Rick Stevens

Large Vision-Language Models (LVLMs) have demonstrated impressive capabilities in multimodal understanding, yet their reasoning abilities remain underexplored. Existing benchmarks tend to focus on perception or text-based comprehension,…

Computation and Language · Computer Science 2025-08-28 Xiang Li , Wenyue Hua , Kaijie Zhu , Lingyao Li , Haoyang Ling , Jinkui Chi , Qi Dou , Jindong Wang , Yongfeng Zhang , Xin Ma , Lizhou Fan

Multimodal Large Language Models (MLLMs) have demonstrated remarkable effectiveness in various general-domain scenarios, such as visual question answering and image captioning. Recently, researchers have increasingly focused on empowering…

Computer Vision and Pattern Recognition · Computer Science 2025-11-04 Yan Shu , Chi Liu , Robin Chen , Derek Li , Bryan Dai

Large language models (LLMs) have performed remarkably well in various natural language processing tasks by benchmarking, including in the Western medical domain. However, the professional evaluation benchmarks for LLMs have yet to be…

Computation and Language · Computer Science 2024-06-04 Wenjing Yue , Xiaoling Wang , Wei Zhu , Ming Guan , Huanran Zheng , Pengfei Wang , Changzhi Sun , Xin Ma

The application of the Multi-modal Large Language Models (MLLMs) in medical clinical scenarios remains underexplored. Previous benchmarks only focus on the capacity of the MLLMs in medical visual question-answering (VQA) or report…

Computation and Language · Computer Science 2024-08-19 Hongcheng Liu , Yusheng Liao , Siqv Ou , Yuhao Wang , Heyang Liu , Yanfeng Wang , Yu Wang

Advancements in Multimodal Large Language Models (MLLMs) have significantly improved medical task performance, such as Visual Question Answering (VQA) and Report Generation (RG). However, the fairness of these models across diverse…

Computer Vision and Pattern Recognition · Computer Science 2024-10-03 Peiran Wu , Che Liu , Canyu Chen , Jun Li , Cosmin I. Bercea , Rossella Arcucci

Large Multi-modality Models (LMMs) have made significant progress in visual understanding and generation, but they still face challenges in General Visual Editing, particularly in following complex instructions, preserving appearance…

Computer Vision and Pattern Recognition · Computer Science 2025-05-28 Xiangyu Zhao , Peiyuan Zhang , Kexian Tang , Xiaorong Zhu , Hao Li , Wenhao Chai , Zicheng Zhang , Renqiu Xia , Guangtao Zhai , Junchi Yan , Hua Yang , Xue Yang , Haodong Duan

Current video benchmarks for multimodal large language models (MLLMs) focus on event recognition, temporal ordering, and long-context recall, but overlook a harder capability required for expert procedural judgment: tracking how ongoing…

Computer Vision and Pattern Recognition · Computer Science 2026-04-13 Xiyang Huang , Jiawei Lin , Keying Wu , Jiaxin Huang , Kailai Yang , Renxiong Wei , Cheng zeng , Jiayi Xiang , Ziyan Kuang , Min Peng , Qianqian Xie , Sophia Ananiadou

Large Vision-Language Models (LVLMs) have made significant strides in the field of video understanding in recent times. Nevertheless, existing video benchmarks predominantly rely on text prompts for evaluation, which often require complex…

Computer Vision and Pattern Recognition · Computer Science 2026-02-04 Yiming Zhao , Yu Zeng , Yukun Qi , YaoYang Liu , Xikun Bao , Lin Chen , Zehui Chen , Qing Miao , Chenxi Liu , Jie Zhao , Feng Zhao

Most organizational data in this world are stored as documents, and visual retrieval plays a crucial role in unlocking the collective intelligence from all these documents. However, existing benchmarks focus on English-only document…

Computer Vision and Pattern Recognition · Computer Science 2025-08-26 Jian Chen , Ming Li , Jihyung Kil , Chenguang Wang , Tong Yu , Ryan Rossi , Tianyi Zhou , Changyou Chen , Ruiyi Zhang

Purpose: Vision-language models (VLMs) have shown promising performance in surgical visual question answering (VQA). However, existing surgical VQA datasets often contain linguistic shortcuts, where question phrasing implicitly constrains…

Computer Vision and Pattern Recognition · Computer Science 2026-05-06 Jongmin Shin , Ka Young Kim , Eunki Cho , Seong Tae Kim , Namkee Oh

Chemical reasoning inherently integrates visual, textual, and symbolic modalities, yet existing benchmarks rarely capture this complexity, often relying on simple image-text pairs with limited chemical semantics. As a result, the actual…

Artificial Intelligence · Computer Science 2025-11-25 Zhiyuan Huang , Baichuan Yang , Zikun He , Yanhong Wu , Fang Hongyu , Zhenhe Liu , Lin Dongsheng , Bing Su

With the rapid growth of large language models (LLMs) and vision-language models (VLMs) in medicine, simply integrating clinical text and medical imaging does not guarantee reliable reasoning. Existing multimodal models often produce…

Artificial Intelligence · Computer Science 2025-12-29 Zelin Zang , Wenyi Gu , Siqi Ma , Dan Yang , Yue Shen , Zhu Zhang , Guohui Fan , Wing-Kuen Ling , Fuji Yang