English
Related papers

Related papers: MRI-Eval: A Tiered Benchmark for Evaluating LLM Pe…

200 papers

The rising prevalence of eye diseases poses a growing public health burden. Large language models (LLMs) offer a promising path to reduce documentation workload and support clinical decision-making. However, few have been tailored for…

Automatic medical report generation has the potential to support clinical diagnosis, reduce the workload of radiologists, and demonstrate potential for enhancing diagnostic consistency. However, current evaluation metrics often fail to…

Computation and Language · Computer Science 2025-08-06 Zhenxuan Zhang , Kinhei Lee , Peiyuan Jing , Weihang Deng , Huichi Zhou , Zihao Jin , Jiahao Huang , Zhifan Gao , Dominic C Marshall , Yingying Fang , Guang Yang

To advance the evaluation of multimodal math reasoning in large multimodal models (LMMs), this paper introduces a novel benchmark, MM-MATH. MM-MATH consists of 5,929 open-ended middle school math problems with visual contexts, with…

Computation and Language · Computer Science 2024-07-03 Kai Sun , Yushi Bai , Ji Qi , Lei Hou , Juanzi Li

Large language models (LLMs) have excelled across domains, also delivering notable performance on the medical evaluation benchmarks, such as MedQA. However, there still exists a significant gap between the reported performance and the…

Computation and Language · Computer Science 2024-06-06 Yuxuan Zhou , Xien Liu , Chen Ning , Ji Wu

Multiple Choice Question (MCQ) answering is a widely used method for evaluating the performance of Large Language Models (LLMs). However, LLMs often exhibit selection bias in MCQ tasks, where their choices are influenced by factors like…

Computation and Language · Computer Science 2025-12-01 Blessed Guda , Lawrence Francis , Gabrial Zencha Ashungafac , Carlee Joe-Wong , Moise Busogi

Evaluating the performance of Grammatical Error Correction (GEC) systems is a challenging task due to its subjectivity. Designing an evaluation metric that is as objective as possible is crucial to the development of GEC task. However,…

Computation and Language · Computer Science 2023-10-18 Jingheng Ye , Yinghui Li , Qingyu Zhou , Yangning Li , Shirong Ma , Hai-Tao Zheng , Ying Shen

State-of-the-art large language models (LLMs) are now claiming remarkable supported context lengths of 256k or even more. In contrast, the average context lengths of mainstream benchmarks are insufficient (5k-21k), and they suffer from…

Computation and Language · Computer Science 2025-10-23 Tao Yuan , Xuefei Ning , Dong Zhou , Zhijie Yang , Shiyao Li , Minghui Zhuang , Zheyue Tan , Zhuyu Yao , Dahua Lin , Boxun Li , Guohao Dai , Shengen Yan , Yu Wang

Reference labels for machine-learning benchmarks are increasingly synthesized with LLM assistance, but their reliability remains underexamined. We audit MedCalc-Bench, a clinical benchmark for medical score computation whose labels were…

Artificial Intelligence · Computer Science 2026-04-14 Junze Ye , Daniel Tawfik , Alex J. Goodell , Nikhil V. Kotha , Mark K. Buyyounouski , Mohsen Bayati

Multimodal Large Language Models (MLLMs) are evaluated on various benchmarks, such as image captioning, visual question answering, and reasoning. However, many of these benchmarks include overly simple or uninformative samples, complicating…

Purpose: We address the challenge of inaccurate parameter estimation in diffusion MRI when the signal-to-noise ratio (SNR) is very low, as in the spinal cord. The accuracy of conventional maximum-likelihood estimation (MLE) depends highly…

Accurate preoperative assessment of lymph node (LN) metastasis in rectal cancer guides treatment decisions, yet conventional MRI evaluation based on morphological criteria shows limited diagnostic performance. While some artificial…

Machine Learning · Computer Science 2025-07-16 Yaoxian Dong , Yifan Gao , Haoyue Li , Yanfen Cui , Xin Gao

Computed Tomography (CT) plays a crucial role in clinical diagnosis, but the growing demand for CT examinations has raised concerns about diagnostic errors. While Multimodal Large Language Models (MLLMs) demonstrate promising comprehension…

Computer Vision and Pattern Recognition · Computer Science 2025-06-25 Sunggu Kyung , Hyungbin Park , Jinyoung Seo , Jimin Sung , Jihyun Kim , Dongyeong Kim , Wooyoung Jo , Yoojin Nam , Sangah Park , Taehee Kwon , Sang Min Lee , Namkug Kim

Clinical problem-solving requires processing of semantic medical knowledge such as illness scripts and numerical medical knowledge of diagnostic tests for evidence-based decision-making. As large language models (LLMs) show promising…

Inpatient medication recommendation requires clinicians to repeatedly select specific medications, doses, and routes as a patient's condition evolves. Existing benchmarks formulate this task as admission-level prediction over coarse drug…

Machine Learning · Computer Science 2026-05-15 Shuhao Chen , Weisen Jiang , Changmiao Wang , Xiaoqing Wu , Xuanren Shi , Yu Zhang , James T. Kwok

Multimodal Large Language Model (MLLM) relies on the powerful LLM to perform multimodal tasks, showing amazing emergent abilities in recent studies, such as writing poems based on an image. However, it is difficult for these case studies to…

Computer Vision and Pattern Recognition · Computer Science 2025-10-27 Chaoyou Fu , Peixian Chen , Yunhang Shen , Yulei Qin , Mengdan Zhang , Xu Lin , Jinrui Yang , Xiawu Zheng , Ke Li , Xing Sun , Yunsheng Wu , Rongrong Ji , Caifeng Shan , Ran He

Objective Structured Clinical Examinations (OSCEs) are widely used to assess medical students' communication skills, but scoring interview-based assessments is time-consuming and potentially subject to human bias. This study explored the…

Computation and Language · Computer Science 2025-05-16 Jadon Geathers , Yann Hicke , Colleen Chan , Niroop Rajashekar , Justin Sewell , Susannah Cornes , Rene F. Kizilcec , Dennis Shung

Magnetic resonance imaging (MRI) quality assessment is crucial for clinical decision-making, yet remains challenging due to data scarcity and protocol variability. Traditional approaches face fundamental trade-offs: signal-based methods…

Computer Vision and Pattern Recognition · Computer Science 2025-09-30 Fankai Jia , Daisong Gan , Zhe Zhang , Zhaochi Wen , Chenchen Dan , Dong Liang , Haifeng Wang

Mastering medical knowledge is crucial for medical-specific LLMs. However, despite the existence of medical benchmarks like MedQA, a unified framework that fully leverages existing knowledge bases to evaluate LLMs' mastery of medical…

Computation and Language · Computer Science 2024-10-03 Yuxuan Zhou , Xien Liu , Chen Ning , Xiao Zhang , Ji Wu

Radiology report evaluation is a crucial part of radiologists' training and plays a key role in ensuring diagnostic accuracy. As part of the standard reporting workflow, a junior radiologist typically prepares a preliminary report, which is…

Computation and Language · Computer Science 2025-10-07 Beth Pearson , Ahmed Adnan , Zahraa S. Abdallah

Scaling pre-training compute has proven effective for achieving mulitlinguality, but does the same hold for test-time scaling? In this work, we introduce MCLM, a multilingual math benchmark featuring competition-level problems in 55…

Computation and Language · Computer Science 2025-08-04 Guijin Son , Jiwoo Hong , Hyunwoo Ko , James Thorne