English
Related papers

Related papers: How to Evaluate Medical AI

200 papers

Benchmarks are pivotal in driving AI progress, and invalid benchmark questions frequently undermine their reliability. Manually identifying and correcting errors among thousands of benchmark questions is not only infeasible but also a…

Prognostic and diagnostic AI-based medical devices hold immense promise for advancing healthcare, yet their rapid development has outpaced the establishment of appropriate validation methods. Existing approaches often fall short in…

Machine Learning · Computer Science 2024-09-10 Florian Hellmeier , Kay Brosien , Carsten Eickhoff , Alexander Meyer

Ensuring safe and effective use of AI requires understanding and anticipating its performance on novel tasks, from advanced scientific challenges to transformed workplace activities. So far, benchmarking has guided progress in AI, but it…

Explainable artificial intelligence (XAI) plays an indispensable role in demystifying the decision-making processes of AI, especially within the healthcare industry. Clinicians rely heavily on detailed reasoning when making a diagnosis,…

Computer Vision and Pattern Recognition · Computer Science 2024-03-26 Anna Stubbin , Thompson Chyrikov , Jim Zhao , Christina Chajo

Screening patients for clinical trial eligibility remains a manual, time-consuming, and resource-intensive process. We present a secure, scalable proof-of-concept system for Artificial Intelligence (AI)-augmented patient-trial matching that…

The automation of AI R&D (AIRDA) could have significant implications, but its extent and ultimate effects remain uncertain. We need empirical data to resolve these uncertainties, but existing data (primarily capability benchmarks) may not…

Computers and Society · Computer Science 2026-03-09 Alan Chan , Ranay Padarath , Joe Kwon , Hilary Greaves , Markus Anderljung

Objective: Breast cancer screening is of great significance in contemporary women's health prevention. The existing machines embedded in the AI system do not reach the accuracy that clinicians hope. How to make intelligent systems more…

Computer Vision and Pattern Recognition · Computer Science 2021-07-13 Jian Dai , Shuge Lei , Licong Dong , Xiaona Lin , Huabin Zhang , Desheng Sun , Kehong Yuan

VLA models have achieved remarkable progress in embodied intelligence; however, their evaluation remains largely confined to simulations or highly constrained real-world settings. This mismatch creates a substantial reality gap, where…

Current medical retrieval benchmarks primarily emphasize lexical or shallow semantic similarity, overlooking the reasoning-intensive demands that are central to clinical decision-making. In practice, physicians often retrieve authoritative…

Information Retrieval · Computer Science 2026-04-07 Xiangxu Zhang , Lei Li , Xiao Zhou , Zheng Liu

What we expect from radiology AI algorithms will shape the selection and implementation of AI in the radiologic practice. In this paper I consider prevailing expectations of AI and compare them to expectations that we have of human readers.…

Computers and Society · Computer Science 2021-05-14 Maciej A. Mazurowski

We introduce a novel framework for incorporating human expertise into algorithmic predictions. Our approach leverages human judgment to distinguish inputs which are algorithmically indistinguishable, or "look the same" to predictive…

Machine Learning · Computer Science 2024-10-31 Rohan Alur , Manish Raghavan , Devavrat Shah

Explainable AI (XAI) is a rapidly growing domain with a myriad of proposed methods as well as metrics aiming to evaluate their efficacy. However, current studies are often of limited scope, examining only a handful of XAI methods and…

Computer Vision and Pattern Recognition · Computer Science 2025-01-03 Lukas Klein , Carsten T. Lüth , Udo Schlegel , Till J. Bungert , Mennatallah El-Assady , Paul F. Jäger

Background: Globally we face a projected shortage of 11 million healthcare practitioners by 2030, and administrative burden consumes 50% of clinical time. Artificial intelligence (AI) has the potential to help alleviate these problems.…

Human-Computer Interaction · Computer Science 2025-08-01 Hashim Hayat , Maksim Kudrautsau , Evgeniy Makarov , Vlad Melnichenko , Tim Tsykunou , Piotr Varaksin , Matt Pavelle , Adam Z. Oskowitz

Deep learning algorithms have shown promising results in visual question answering (VQA) tasks, but a more careful look reveals that they often do not understand the rich signal they are being fed with. To understand and better measure the…

Computer Vision and Pattern Recognition · Computer Science 2021-09-20 Daniel Rosenberg , Itai Gat , Amir Feder , Roi Reichart

Artificial intelligence (AI) holds great promise for supporting clinical trials, from patient recruitment and endpoint assessment to treatment response prediction. However, deploying AI without safeguards poses significant risks,…

Machine Learning · Computer Science 2025-10-09 Yao Chen , David Ohlssen , Aimee Readie , Gregory Ligozio , Ruvie Martin , Thibaud Coroller

The integration of artificial intelligence (AI), particularly Convolutional Neural Networks (CNNs), into dermatological diagnosis demonstrates substantial clinical potential. While existing literature predominantly benchmarks algorithmic…

Computer Vision and Pattern Recognition · Computer Science 2026-04-02 Loris Cino , Pier Luigi Mazzeo , Alessandro Martella , Giulia Radi , Renato Rossi , Cosimo Distante

Trustworthiness and transparency are essential for the clinical adoption of artificial intelligence (AI) in healthcare and biomedical research. Recent deep research systems aim to accelerate evidence-grounded scientific discovery by…

With the proliferation of large language models (LLMs) in the medical domain, there is increasing demand for improved evaluation techniques to assess their capabilities. However, traditional metrics like F1 and ROUGE, which rely on token…

Computation and Language · Computer Science 2025-05-20 Xiechi Zhang , Zetian Ouyang , Linlin Wang , Gerard de Melo , Zhu Cao , Xiaoling Wang , Ya Zhang , Yanfeng Wang , Liang He

Human evaluations play a central role in training and assessing AI models, yet these data are rarely treated as measurements subject to systematic error. This paper integrates psychometric rater models into the AI pipeline to improve the…

Artificial Intelligence · Computer Science 2026-02-27 Jodi M. Casabianca , Maggie Beiting-Parrish
‹ Prev 1 4 5 6 7 8 10 Next ›