English
Related papers

Related papers: Mind the Performance Gap: Examining Dataset Shift …

200 papers

Social contexts -- such as families, schools, and neighborhoods -- shape life outcomes. The key question is not simply whether they matter, but rather for whom and under what conditions. Here, we argue that prediction gaps -- differences in…

Social and Information Networks · Computer Science 2025-07-01 Javier Garcia-Bernardo , Eva Jaspers , Weverthon Machado , Samuel Plach , Erik Jan van Leeuwen

Estimating the test performance of software AI-based medical devices under distribution shifts is crucial for evaluating the safety, efficiency, and usability prior to clinical deployment. Due to the nature of regulated medical device…

Machine Learning · Computer Science 2022-07-14 Charles Lu , Syed Rakin Ahmed , Praveer Singh , Jayashree Kalpathy-Cramer

Accuracy and interpretability are two dominant features of successful predictive models. Typically, a choice must be made in favor of complex black box models such as recurrent neural networks (RNN) for accuracy versus less accurate but…

Machine Learning · Computer Science 2017-02-28 Edward Choi , Mohammad Taha Bahadori , Joshua A. Kulas , Andy Schuetz , Walter F. Stewart , Jimeng Sun

In clinical trials studying paired parts of a subject with binary outcomes, it is expected to collect measurements bilaterally. However, there are cases where subjects contribute measurements for only one part. By utilizing combined data,…

Applications · Statistics 2024-03-06 Shuyi Liang , Kai-Tai Fang , Xin-Wei Huang , Yijing Xin , Chang-Xing Ma

Predictive models for clinical outcomes that are accurate on average in a patient population may underperform drastically for some subpopulations, potentially introducing or reinforcing inequities in care access and quality. Model training…

Machine Learning · Statistics 2022-02-03 Stephen R. Pfohl , Haoran Zhang , Yizhe Xu , Agata Foryciarz , Marzyeh Ghassemi , Nigam H. Shah

Recent work has shown that the performance of machine learning models can vary substantially when models are evaluated on data drawn from a distribution that is close to but different from the training distribution. As a result, predicting…

Machine Learning · Computer Science 2021-08-23 Devin Guillory , Vaishaal Shankar , Sayna Ebrahimi , Trevor Darrell , Ludwig Schmidt

In a stratified clinical trial design with time to event end points, stratification factors are often accounted for the log-rank test and the Cox regression analyses. In this work, we have evaluated the impact of inclusion of stratification…

Applications · Statistics 2020-06-30 Madan G. Kundu , Shoubhik Mondal

Performance regression testing is essential in large-scale continuous-integration (CI) systems, yet executing full performance suites for every commit is prohibitively expensive. Prior work on performance regression prediction and batch…

Software Engineering · Computer Science 2026-04-02 Ali Sayedsalehi , Peter C. Rigby , Gregory Mierzwinski

Evaluating and validating the performance of prediction models is a fundamental task in statistics, machine learning, and their diverse applications. However, developing robust performance metrics for competing risks time-to-event data…

Methodology · Statistics 2025-07-22 Zian Zhuang , Wen Su , Eric Kawaguchi , Gang Li

Interest in targeted disease prevention has stimulated development of models that assign risks to individuals, using their personal covariates. We need to evaluate these models, and to quantify the gains achieved by expanding a model with…

Methodology · Statistics 2009-06-16 Alice S. Whittemore

COVID-19 has challenged health systems to learn how to learn. This paper describes the context, methods and challenges for learning to improve COVID-19 care at one academic health center. Challenges to learning include: (1) choosing a right…

Objectives: To evaluate the consequences of the framing of machine learning risk prediction models. We evaluate how framing affects model performance and model learning in four different approaches previously applied in published…

In the era of increasingly complex AI models for time series forecasting, progress is often measured by marginal improvements on benchmark leaderboards. However, this approach suffers from a fundamental flaw: standard evaluation metrics…

Machine Learning · Computer Science 2026-05-28 Wanjin Feng , Yuan Yuan , Jingtao Ding , Yong Li

Predicting relative risk (RR) of spatial clusters is a complex task in public health that can be achieved through various statistical and machine-learning methods for different time intervals. However, high-resolution longitudinal data is…

Methodology · Statistics 2025-12-23 Lyza Iamrache , Kamel Rekab , Majid Bani-Yagoub , Julia Pluta , Abdelghani Mehailia

Clinical trials provide essential guidance for practicing Evidence-Based Medicine, though often accompanying with unendurable costs and risks. To optimize the design of clinical trials, we introduce a novel Clinical Trial Result Prediction…

Computation and Language · Computer Science 2020-10-13 Qiao Jin , Chuanqi Tan , Mosha Chen , Xiaozhong Liu , Songfang Huang

In program evaluations, units can often anticipate the implementation of a new policy before it occurs. Such anticipatory behavior can lead to units' outcomes becoming dependent on their future treatment assignments. In this paper, I employ…

Econometrics · Economics 2022-12-02 Aibo Gong

Propensity score weighting approaches have been widely implemented in clinical research to estimate the effects of a treatment or exposure while mitigating the risk of confounding in the absence of random assignment. In practice, when…

Methodology · Statistics 2026-04-17 Emma K. Mackay , Amol A. Verma , Fahad Razak , Surain B. Roberts

X-ray and computed tomography (CT) scanning technologies for COVID-19 screening have gained significant traction in AI research since the start of the coronavirus pandemic. Despite these continuous advancements for COVID-19 screening, many…

Image and Video Processing · Electrical Eng. & Systems 2020-05-06 Brian D Goodwin , Corey Jaskolski , Can Zhong , Herick Asmani

Predictions under interventions are estimates of what a person's risk of an outcome would be if they were to follow a particular treatment strategy, given their individual characteristics. Such predictions can give important input to…

Methodology · Statistics 2025-06-17 Ruth H. Keogh , Nan van Geloven

Respiratory failure is the one of major causes of death in critical care unit. During the outbreak of COVID-19, critical care units experienced an extreme shortage of mechanical ventilation because of respiratory failure related syndromes.…

Machine Learning · Computer Science 2021-09-08 Yilin Yin , Chun-An Chou