English
Related papers

Related papers: How to Evaluate Medical AI

200 papers

Diagnostic errors in healthcare persist as a critical challenge, with increasing numbers of patients turning to online resources for health information. While AI-powered healthcare chatbots show promise, there exists no standardized and…

Artificial Intelligence · Computer Science 2024-12-18 Deep Bhatt , Surya Ayyagari , Anuruddh Mishra

Radiology reports for the same patient examination may contain clinically meaningful discrepancies arising from interpretation differences, reporting variability, or evolving assessments. Systematic analysis of such discrepancies is…

Computer Vision and Pattern Recognition · Computer Science 2026-03-10 Zhaoyi Sun , Minal Jagtiani , Wen-wai Yim , Fei Xia , Martin Gunn , Meliha Yetisgen , Asma Ben Abacha

Medicine is inherently pluralistic. Principles such as autonomy, beneficence, nonmaleficence, and justice routinely conflict, and such ethical dilemmas often sharply divide reasonable physicians. Good clinical practice navigates these…

Large language models (LLMs) are gaining increasing interests to improve clinical efficiency for medical diagnosis, owing to their unprecedented performance in modelling natural language. Ensuring the safe and reliable clinical…

Computation and Language · Computer Science 2024-03-26 Lei Liu , Xiaoyan Yang , Fangzhou Li , Chenfei Chi , Yue Shen , Shiwei Lyu Ming Zhang , Xiaowei Ma , Xiangguo Lyu , Liya Ma , Zhiqiang Zhang , Wei Xue , Yiran Huang , Jinjie Gu

Artificial intelligence (AI) has become increasingly central to precision medicine by enabling the integration and interpretation of multimodal data, yet implementation in clinical settings remains limited. This paper provides a scoping…

Artificial Intelligence · Computer Science 2025-10-17 Göktuğ Bender , Samer Faraj , Anand Bhardwaj

While the capabilities and utility of AI systems have advanced, rigorous norms for evaluating these systems have lagged. Grand claims, such as models achieving general reasoning capabilities, are supported with model performance on narrow…

In recent years, AI systems in the medical domain have advanced significantly. However, despite outperforming humans, they are rarely used in practice since it is often not clear how they make their decisions. Optimal explanation and…

Human-Computer Interaction · Computer Science 2026-05-06 Julia Wagner , Tim Schlippe

Adaptive interventions (AIs) are increasingly becoming popular in medical and behavioral sciences. An AI is a sequence of individualized intervention options that specify for whom and under what conditions different intervention options…

Applications · Statistics 2018-12-18 Palash Ghosh , Inbal Nahum-Shani , Bonnie Spring , Bibhas Chakraborty

Accurate interpretation of lab results is crucial in clinical medicine, yet most patient portals use universal normal ranges, ignoring conditional factors like age and gender. This study introduces Lab-AI, an interactive system that offers…

Computation and Language · Computer Science 2025-04-25 Xiaoyu Wang , Haoyong Ouyang , Balu Bhasuran , Xiao Luo , Karim Hanna , Mia Liza A. Lustria , Carl Yang , Zhe He

Large artificial intelligence (AI) models have garnered significant attention for their remarkable, often "superhuman", performance on standardized benchmarks. However, when these models are deployed in high-stakes verticals such as…

Artificial Intelligence · Computer Science 2025-09-26 Gaurav Verma , Jiawei Zhou , Mohit Chandra , Srijan Kumar , Munmun De Choudhury

Reliable clinical decision support requires medical AI agents capable of safe, multi-step reasoning over structured electronic health records (EHRs). While large language models (LLMs) show promise in healthcare, existing benchmarks…

Artificial Intelligence · Computer Science 2026-01-15 Ananya Mantravadi , Shivali Dalmia , Abhishek Mukherji

Artificial intelligence has demonstrated significant potential in clinical decision-making; however, developing models capable of adapting to diverse real-world scenarios and performing complex diagnostic reasoning remains a major…

Computer Vision and Pattern Recognition · Computer Science 2025-08-18 Ronghao Xu , Zhen Huang , Yangbo Wei , Xiaoqian Zhou , Zikang Xu , Ting Liu , Zihang Jiang , S. Kevin Zhou

In the precision medicine era, there is a growing need for precision radiotherapy where the planned radiation dose needs to be optimally determined by considering a myriad of patient-specific information in order to ensure treatment…

Machine Learning · Statistics 2022-02-10 Wenbo Sun , Dipesh Niraula , Issam El Naqa , Randall K Ten Haken , Ivo D Dinov , Kyle Cuneo , Judy Jin

Artificial Intelligence is increasingly introduced into systems engineering activities, particularly within requirements engineering, where quality assessment and validation remain heavily dependent on expert judgment. While recent AI tools…

Software Engineering · Computer Science 2026-04-17 Oz Levy , Ilya Dikman , Natan Levy , Michael Winokur

Artificial intelligence (AI)-based methods are showing promise in multiple medical-imaging applications. Thus, there is substantial interest in clinical translation of these methods, requiring in turn, that they be evaluated rigorously. In…

The development of Artificial Intelligence (AI), including AI in Science (AIS), should be done following the principles of responsible AI. Progress in responsible AI is often quantified through evaluation metrics, yet there has been less…

Computers and Society · Computer Science 2025-10-31 Theresia Veronika Rampisela , Maria Maistro , Tuukka Ruotsalo , Christina Lioma

Background: Evaluating AI-generated treatment plans is a key challenge as AI expands beyond diagnostics, especially with new reasoning models. This study compares plans from human experts and two AI models (a generalist and a reasoner),…

Artificial Intelligence · Computer Science 2025-07-09 Dipayan Sengupta , Saumya Panda

Medical data poses a daunting challenge for AI algorithms: it exists in many different modalities, experiences frequent distribution shifts, and suffers from a scarcity of examples and labels. Recent advances, including transformers and…

Scientific discovery begins with ideas, yet evaluating early-stage research concepts is a subtle and subjective human judgment. As large language models (LLMs) are increasingly tasked with generating scientific hypotheses, most systems…

Human-Computer Interaction · Computer Science 2026-03-26 Lingyu Zhang , Mitchell Wang , Boyuan Chen

Recent Artificial Intelligence (AI) models have matched or exceeded human experts in several benchmarks of biomedical task performance, but surgical benchmarks in particular are often missing from prominent medical benchmark suites. Since…