English
Related papers

Related papers: Docs are ROCs: A simple off-the-shelf approach for…

200 papers

As the most important tool to provide high-level evidence-based medicine, researchers can statistically summarize and combine data from multiple studies by conducting meta-analysis. In meta-analysis, mean differences are frequently used…

Methodology · Statistics 2018-01-30 Dehui Luo , Xiang Wan , Jiming Liu , Tiejun Tong

The area under the curve (AUC) of summary receiver operating characteristic (SROC) curve is a primary statistical outcome for meta-analysis of diagnostic test accuracy studies (DTA). However, its confidence interval has not been reported in…

Applications · Statistics 2022-08-02 Hisashi Noma , Yuki Matsushima , Ryota Ishii

Diagnostic tests are of critical importance in health care and medical research. Motivated by the impact that atypical and outlying test outcomes might have on the assessment of the discriminatory ability of a diagnostic test, we develop a…

We have carried out a pilot study on a standard collection of electrocardiograms from patients who suffer from congestive heart failure, and subjects without cardiac pathology, using receiver-operating-characteristic (ROC) analysis. The…

chao-dyn · Physics 2007-05-23 Stefan Thurner , Markus C. Feurstein , Malvin C. Teich

Comparing model performances on benchmark datasets is an integral part of measuring and driving progress in artificial intelligence. A model's performance on a benchmark dataset is commonly assessed based on a single or a small set of…

Artificial Intelligence · Computer Science 2021-11-09 Kathrin Blagec , Georg Dorffner , Milad Moradi , Matthias Samwald

In modern dynamic constantly developing society, more and more people suffer from chronic and serious diseases and doctors and patients need special and sophisticated medical and health support. Accordingly, prominent health stakeholders…

Computers and Society · Computer Science 2022-08-10 Mirjana Ivanovic , Serge Autexier , Miltiadis Kokkonidis

We consider the problem of estimating a dose-response curve. Continuous treatments arise often in practice, e.g. in the form of time spent on an operation, distance traveled to a location or dosage of a drug. Letting $A$ denote a continuous…

Methodology · Statistics 2026-04-14 Matteo Bonvini , Edward H. Kennedy

In human-AI collaboration systems for critical applications, in order to ensure minimal error, users should set an operating point based on model confidence to determine when the decision should be delegated to human experts. Samples for…

Artificial Intelligence · Computer Science 2023-10-13 Sara Sangalli , Ertunc Erdil , Ender Konukoglu

The ROC (receiver operating characteristic) curve is a widely used device for assessing decision-making systems. It seems surprising, in view of its history dating back to World War Two, that the assignment of uncertainties to a ROC curve…

Data Analysis, Statistics and Probability · Physics 2024-08-19 M. P. Fewell

To assess the classification accuracy of a continuous diagnostic result, the receiver operating characteristic (ROC) curve is commonly used in applications. The partial area under the ROC curve (pAUC) is one of widely accepted summary…

Applications · Statistics 2011-03-11 Hung Hung , Chin-Tsang Chiang

We present results from a pilot experiment to measure if machine recommendations can debias human perceptual biases in visualization tasks. We specifically studied the ``pull-down'' effect, i.e., people underestimate the average position of…

Human-Computer Interaction · Computer Science 2023-11-03 Ross Geuy , Nate Rising , Tiancheng Shi , Meng Ling , Jian Chen

Artificial intelligence (AI) systems are deployed as collaborators in human decision-making. Yet, evaluation practices focus primarily on model accuracy rather than whether human-AI teams are prepared to collaborate safely and effectively.…

Human-Computer Interaction · Computer Science 2026-03-20 Min Hun Lee

The area under the ROC curve is widely used as a measure of performance of classification rules. However, it has recently been shown that the measure is fundamentally incoherent, in the sense that it treats the relative severities of…

Methodology · Statistics 2013-08-02 David J. Hand , Christoforos Anagnostopoulos

AI models are increasingly deployed in live clinical environments where they must perform reliably across complex, high-stakes workflows that standard training and validation datasets were never designed to capture. Evaluating these systems…

Artificial Intelligence · Computer Science 2026-05-12 Prasanna Desikan , Harshit Rajgarhia , Shivali Dalmia , Ananya Mantravadi

When technical requirements are high, and patient outcomes are critical, opportunities for monitoring and improving surgical skills via objective motion analysis feedback may be particularly beneficial. This narrative review synthesises…

Computer Vision and Pattern Recognition · Computer Science 2024-02-20 Merryn D. Constable , Hubert P. H. Shum , Stephen Clark

Two main approaches for evaluating the quality of machine-generated rationales are: 1) using human rationales as a gold standard; and 2) automated metrics based on how rationales affect model behavior. An open question, however, is how…

Computation and Language · Computer Science 2020-10-13 Samuel Carton , Anirudh Rathore , Chenhao Tan

While automated driving is often advertised with better-than-human driving performance, this work reviews that it is nearly impossible to provide direct statistical evidence on the system level that this is actually the case. The amount of…

Machine Learning · Computer Science 2021-12-10 Hanno Gottschalk , Matthias Rottmann , Maida Saltagic

Most clinical AI systems operate as prediction engines -- producing labels or risk scores -- yet real clinical reasoning is a time-bounded, sequential control problem under uncertainty. Clinicians interleave information gathering with…

Artificial Intelligence · Computer Science 2026-01-21 Dipayan Sengupta , Saumya Panda

Most binary classifiers work by processing the input to produce a scalar response and comparing it to a threshold value. The various measures of classifier performance assume, explicitly or implicitly, probability distributions $P_s$ and…

Machine Learning · Computer Science 2019-09-24 Luma Omar , Ioannis Ivrissimtzis

We introduce a novel framework for incorporating human expertise into algorithmic predictions. Our approach leverages human judgment to distinguish inputs which are algorithmically indistinguishable, or "look the same" to predictive…

Machine Learning · Computer Science 2024-10-31 Rohan Alur , Manish Raghavan , Devavrat Shah
‹ Prev 1 3 4 5 6 7 10 Next ›