中文
相关论文

相关论文: Docs are ROCs: A simple off-the-shelf approach for…

200 篇论文

Randomized clinical trials are the gold standard when estimating the average treatment effect. However, they are usually not a random sample from the real-world population because of the inclusion/exclusion rules. Meanwhile, observational…

统计方法学 · 统计学 2024-12-11 Kuan Jiang , Wenjie Hu , Shu Yang , Xinxing Lai , Xiaohua Zhou

Estimating what would be an individual's potential response to varying levels of exposure to a treatment is of high practical relevance for several important fields, such as healthcare, economics and public policy. However, existing methods…

机器学习 · 计算机科学 2020-12-11 Patrick Schwab , Lorenz Linhardt , Stefan Bauer , Joachim M. Buhmann , Walter Karlen

Purpose: Artificial intelligence (AI) solutions for medical diagnosis require thorough evaluation to demonstrate that performance is maintained for all patient sub-groups and to ensure that proposed improvements in care will be delivered…

Aim: provide a methodological framework for the process of clinical tests, clinical acceptance, and scientific assessment of algorithms and software based on the artificial intelligence (AI) technologies. Clinical tests are considered as a…

Accurate quantification of model uncertainty has long been recognized as a fundamental requirement for trusted AI. In regression tasks, uncertainty is typically quantified using prediction intervals calibrated to a specific operating point,…

机器学习 · 计算机科学 2021-06-03 Jiri Navratil , Benjamin Elder , Matthew Arnold , Soumya Ghosh , Prasanna Sattigeri

Clinical trials usually target average treatment effects, but treatment decisions are made for individuals. This tension motivates a common criticism of evidence-based medicine: a treatment that is beneficial on average may be inappropriate…

应用统计 · 统计学 2026-05-29 Zach Shahn , Mats Stensrud

Current medical AI systems are often limited to narrow applications, hindering widespread adoption. We present MedVersa, a generalist foundation model trained on tens of millions of compiled medical instances. MedVersa unlocks generalist…

计算机视觉与模式识别 · 计算机科学 2025-06-11 Hong-Yu Zhou , Julián Nicolás Acosta , Subathra Adithan , Suvrankar Datta , Eric J. Topol , Pranav Rajpurkar

Receiver operating characteristic (ROC) curve is an informative tool in binary classification and Area Under ROC Curve (AUC) is a popular metric for reporting performance of binary classifiers. In this paper, first we present a…

机器学习 · 计算机科学 2021-09-14 Khashayar Namdar , Masoom A. Haider , Farzad Khalvati

The area under the receiver-operating characteristic curve (AUC) has become a popular index not only for measuring the overall prediction capacity of a marker but also the association strength between continuous and binary variables. In the…

统计方法学 · 统计学 2023-10-12 Pablo Martinez-Camblor , Sonia Perez-Fernandez , Lucas L. Dwiel , Wilder T. Doucette

In analysis of binary outcomes, the receiver operator characteristic (ROC) curve is heavily used to show the performance of a model or algorithm. The ROC curve is informative about the performance over a series of thresholds and can be…

统计计算 · 统计学 2020-08-10 John Muschelli

Language models (LMs) represent an emerging paradigm within artificial intelligence, with applications throughout the medical enterprise. A comprehensive understanding of the clinical task and awareness of the variability in performance…

机器学习 · 计算机科学 2026-03-09 Victor Garcia , Mariia Sidulova , Aldo Badano

A key trait of stochastic optimizers is that multiple runs of the same optimizer in attempting to solve the same problem can produce different results. As a result, their performance is evaluated over several repeats, or runs, on the…

机器学习 · 计算机科学 2026-05-18 Moslem Noori , Elisabetta Valiante , Thomas Van Vaerenbergh , Masoud Mohseni , Ignacio Rozada

AI-driven clinical text classification is vital for explainable automated retrieval of population-level health information. This work investigates whether human-based clinical rationales can serve as additional supervision to improve both…

计算与语言 · 计算机科学 2025-07-30 Christoph Metzner , Shang Gao , Drahomira Herrmannova , Heidi A. Hanson

Performance comparisons are fundamental in medical imaging Artificial Intelligence (AI) research, often driving claims of superiority based on relative improvements in common performance metrics. However, such claims frequently rely solely…

Estimating the performance of a machine learning system is a longstanding challenge in artificial intelligence research. Today, this challenge is especially relevant given the emergence of systems which appear to increasingly outperform…

机器学习 · 计算机科学 2021-09-17 Qiongkai Xu , Christian Walder , Chenchen Xu

If artificial intelligence (AI) is to be applied in safety-critical domains, its performance needs to be evaluated reliably. The present study aimed to understand how humans evaluate AI systems for person detection in automatic train…

人机交互 · 计算机科学 2025-04-04 Romy Müller

Human-AI collaboration for decision-making strives to achieve team performance that exceeds the performance of humans or AI alone. However, many factors can impact success of Human-AI teams, including a user's domain expertise, mental…

The receiver operating characteristic (ROC) curve is the most popular tool used to evaluate the discriminatory capability of diagnostic tests/biomarkers measured on a continuous scale when distinguishing between two alternative disease…

统计方法学 · 统计学 2021-03-22 Maria Xose Rodriguez-Alvarez , Vanda Inacio

In machine learning, we traditionally evaluate the performance of a single model, averaged over a collection of test inputs. In this work, we propose a new approach: we measure the performance of a collection of models when evaluated on a…

机器学习 · 计算机科学 2022-06-08 Gal Kaplun , Nikhil Ghosh , Saurabh Garg , Boaz Barak , Preetum Nakkiran

Background: Medical decision-making impacts both individual and public health. Clinical scores are commonly used among a wide variety of decision-making models for determining the degree of disease deterioration at the bedside. AutoScore…