English
Related papers

Related papers: LiveMedBench: A Contamination-Free Medical Benchma…

200 papers

While Large Language Models (LLMs) excel on standardized medical exams, high scores often fail to translate to high-quality responses for real-world medical queries. Current evaluations rely heavily on multiple-choice questions, failing to…

Large Language Models (LLMs) applied to code-related applications have emerged as a prominent field, attracting significant interest from both academia and industry. However, as new and improved LLMs are developed, existing evaluation…

Software Engineering · Computer Science 2024-06-07 Naman Jain , King Han , Alex Gu , Wen-Ding Li , Fanjia Yan , Tianjun Zhang , Sida Wang , Armando Solar-Lezama , Koushik Sen , Ion Stoica

Evaluating large language models (LLMs) in medicine is crucial because medical applications require high accuracy with little room for error. Current medical benchmarks have three main types: medical exam-based, comprehensive medical, and…

With the proliferation of Large Language Models (LLMs) in diverse domains, there is a particular need for unified evaluation standards in clinical medical scenarios, where models need to be examined very thoroughly. We present CliMedBench,…

Computation and Language · Computer Science 2024-10-07 Zetian Ouyang , Yishuai Qiu , Linlin Wang , Gerard de Melo , Ya Zhang , Yanfeng Wang , Liang He

Test set contamination, wherein test data from a benchmark ends up in a newer model's training set, is a well-documented obstacle for fair LLM evaluation and can quickly render benchmarks obsolete. To mitigate this, many recent benchmarks…

Large language models (LLMs) show significant potential in healthcare, prompting numerous benchmarks to evaluate their capabilities. However, concerns persist regarding the reliability of these benchmarks, which often lack clinical…

Computation and Language · Computer Science 2026-04-30 Wenting Chen , Guo Yu , Yiu-Fai Cheung , Meidan Ding , Jie Liu , Zizhan Ma , Wenxuan Wang , Linlin Shen

The reliability of medical LLM evaluation is critically undermined by data contamination and knowledge obsolescence, leading to inflated scores on static benchmarks. To address these challenges, we introduce LiveClin, a live benchmark…

Machine Learning · Computer Science 2026-02-20 Xidong Wang , Shuqi Guo , Yue Shen , Junying Chen , Jian Wang , Jinjie Gu , Ping Zhang , Lei Liu , Benyou Wang

The emergence of various medical large language models (LLMs) in the medical domain has highlighted the need for unified evaluation standards, as manual evaluation of LLMs proves to be time-consuming and labor-intensive. To address this…

Computation and Language · Computer Science 2023-12-21 Yan Cai , Linlin Wang , Ye Wang , Gerard de Melo , Ya Zhang , Yanfeng Wang , Liang He

In contrast to their remarkable performance on general knowledge QA, the true abilities of Large Language Models (LLMs) in tasks demanding deep, specialized reasoning, such as in protein biology, have yet to be thoroughly investigated.…

Quantitative Methods · Quantitative Biology 2025-12-30 Dingyi Rong , Zijian Chen , Qi Jia , Kaiwei Zhang , Haotian Lu , Guangtao Zhai , Ning Liu

Clinical reasoning in medicine is a hypothesis-driven process where physicians refine diagnoses from limited information through targeted history, physical examination, and diagnostic investigations. In contrast, current medical benchmarks…

Machine Learning · Computer Science 2025-10-14 Christopher Chiu , Silviu Pitis , Mihaela van der Schaar

Large Language Models (LLMs) are increasingly used for clinical decision support, where hallucinations and unsafe suggestions may pose direct risks to patient safety. These risks are hard to assess: subtle clinical errors are often missed…

Computation and Language · Computer Science 2026-05-14 Yinzhu Chen , Abdine Maiga , Hossein A. Rahmani , Emine Yilmaz

Ensuring the general efficacy and goodness for human beings from medical large language models (LLM) before real-world deployment is crucial. However, a widely accepted and accessible evaluation process for medical LLM, especially in the…

As opposed to evaluating computation and logic-based reasoning, current benchmarks for evaluating large language models (LLMs) in medicine are primarily focused on question-answering involving domain knowledge and descriptive reasoning.…

Large Language Models (LLMs) have demonstrated impressive capabilities across various specialist domains and have been integrated into high-stakes areas such as medicine. However, as existing medical-related benchmarks rarely stress-test…

Computation and Language · Computer Science 2026-03-26 Lin Yang , Yuancheng Yang , Xu Wang , Changkun Liu , Haihua Yang

The adoption of large language models (LLMs) to assist clinicians has attracted remarkable attention. Existing works mainly adopt the close-ended question-answering (QA) task with answer options for evaluation. However, many clinical…

As large language models (LLMs) become increasingly integrated into clinical decision-making, ensuring transparent and trustworthy reasoning is essential. However, existing evaluation strategies of LLMs' medical reasoning capability either…

Multimodal large language models (MLLMs) hold significant potential in medical applications, including disease diagnosis and clinical decision-making. However, these tasks require highly accurate, context-sensitive, and professionally…

Computation and Language · Computer Science 2025-09-01 Meidan Ding , Jipeng Zhang , Wenxuan Wang , Cheng-Yi Li , Wei-Chieh Fang , Hsin-Yu Wu , Haiqin Zhong , Wenting Chen , Linlin Shen

Inaccuracies in existing or generated clinical text may lead to serious adverse consequences, especially if it is a misdiagnosis or incorrect treatment suggestion. With Large Language Models (LLMs) increasingly being used across diverse…

Computation and Language · Computer Science 2026-02-06 Congbo Ma , Yichun Zhang , Yousef Al-Jazzazi , Ahamed Foisal , Laasya Sharma , Yousra Sadqi , Khaled Saleh , Jihad Mallat , Farah E. Shamout

Medical conversational AI (AI) plays a pivotal role in the development of safer and more effective medical dialogue systems. However, existing benchmarks and evaluation frameworks for assessing the information-gathering and diagnostic…

Computation and Language · Computer Science 2026-01-08 Lecheng Gong , Weimin Fang , Ting Yang , Dongjie Tao , Chunxiao Guo , Peng Wei , Bo Xie , Jinqun Guan , Zixiao Chen , Fang Shi , Jinjie Gu , Junwei Liu

As Large Language Models (LLMs) exhibit plateauing performance on conventional benchmarks, a pivotal challenge persists: evaluating their proficiency in complex, open-ended tasks characterizing genuine expert-level cognition. Existing…

‹ Prev 1 2 3 10 Next ›