English
Related papers

Related papers: Performance metrics for intervention-triggering pr…

200 papers

How should researchers analyze randomized experiments in which the main outcome is latent and measured in multiple ways but each measure contains some degree of error? We first identify a critical study-specific noncomparability problem in…

Econometrics · Economics 2026-01-13 Jiawei Fu , Donald P. Green

A primary difficulty with unsupervised discovery of structure in large data sets is a lack of quantitative evaluation criteria. In this work, we propose and investigate several metrics for evaluating and comparing generative models of…

Machine Learning · Computer Science 2020-07-27 Daniel Jiwoong Im , Iljung Kwak , Kristin Branson

Intelligent agents rely on AI/ML functionalities to predict the consequence of possible actions and optimise the policy. However, the effort of the research community in addressing prediction accuracy has been so intense (and successful)…

Machine Learning · Computer Science 2023-10-04 Gianluca Bontempi

Model multiplicity refers to the existence of multiple machine learning models that describe the data equally well but may produce different predictions on individual samples. In medicine, these models can admit conflicting predictions for…

Automated decision systems (ADS) are broadly deployed to inform and support human decision-making across a wide range of consequential settings. However, various context-specific details complicate the goal of establishing meaningful…

Computers and Society · Computer Science 2026-02-05 Inioluwa Deborah Raji , Lydia Liu

Predictions about people, such as their expected educational achievement or their credit risk, can be performative and shape the outcome that they aim to predict. Understanding the causal effect of these predictions on the eventual outcomes…

Machine Learning · Statistics 2022-10-19 Celestine Mendler-Dünner , Frances Ding , Yixin Wang

Presenting a predictive model's performance is a communication bottleneck that threatens collaborations between data scientists and subject matter experts. Accuracy and error metrics alone fail to tell the whole story of a model - its…

Human-Computer Interaction · Computer Science 2025-03-19 Ashley Suh , Gabriel Appleby , Erik W. Anderson , Luca Finelli , Remco Chang , Dylan Cashman

Prescriptive process monitoring approaches leverage historical data to prescribe runtime interventions that will likely prevent negative case outcomes or improve a process's performance. A centerpiece of a prescriptive process monitoring…

Artificial Intelligence · Computer Science 2022-06-17 Mahmoud Shoush , Marlon Dumas

Most machine learning techniques are based upon statistical learning theory, often simplified for the sake of computing speed. This paper is focused on the uncertainty aspect of mathematical modeling in machine learning. Regression analysis…

Machine Learning · Computer Science 2022-06-07 Valentin Arkov

A rich set of frequentist model averaging methods has been developed, but their applications have largely been limited to point prediction, as measuring prediction uncertainty in general settings remains an open problem. In this paper we…

Econometrics · Economics 2025-10-21 Zhongjun Qu , Wendun Wang , Xiaomeng Zhang

Risk prediction models are often advertised as deterministic functions that map covariates to predicted risks. However, they are typically trained using finite samples, and as such, their predictions are inherently uncertain. This…

Methodology · Statistics 2025-06-03 Abdollah Safari , Paul Gustafson , Mohsen Sadatsafavi

Test data measured by medical instruments often carry imprecise ranges that include the true values. The latter are not obtainable in virtually all cases. Most learning algorithms, however, carry out arithmetical calculations that are…

Machine Learning · Computer Science 2020-07-27 Mei Wang , Jianwen Su , Haiqin Lu

In deep learning applications, robustness measures the ability of neural models that handle slight changes in input data, which could lead to potential safety hazards, especially in safety-critical applications. Pre-deployment assessment of…

Software Engineering · Computer Science 2024-04-26 Wenchuan Mu , Kwan Hui Lim

Classification systems are evaluated in a countless number of papers. However, we find that evaluation practice is often nebulous. Frequently, metrics are selected without arguments, and blurry terminology invites misconceptions. For…

Machine Learning · Computer Science 2024-07-03 Juri Opitz

Conventional meta analysis of model performance conducted using datasources from different underlying populations often result in estimates that cannot be interpreted in the context of a well defined target population. In this manuscript we…

Methodology · Statistics 2024-09-23 Jon A. Steingrimsson , Lan Wen , Sarah Voter , Issa J. Dahabreh

Who should we prioritize for intervention when we cannot estimate intervention effects? In many applied domains (e.g., advertising, customer retention, and behavioral nudging) prioritization is guided by predictive models that estimate…

Machine Learning · Computer Science 2025-04-07 Carlos Fernández-Loría , Jorge Loría

Machine learning (ML) models have been quite successful in predicting outcomes in many applications. However, in some cases, domain experts might have a judgment about the expected outcome that might conflict with the prediction of ML…

Machine Learning · Computer Science 2023-05-02 Hogun Park , Aly Megahed , Peifeng Yin , Yuya Ong , Pravar Mahajan , Pei Guo

Detecting other agents and forecasting their behavior is an integral part of the modern robotic autonomy stack, especially in safety-critical scenarios entailing human-robot interaction such as autonomous driving. Due to the importance of…

Robotics · Computer Science 2021-10-08 Boris Ivanovic , Marco Pavone

We consider the estimation of measures of model performance in a target population when covariate and outcome data are available on a sample from some source population and covariate data, but not outcome data, are available on a simple…

Methodology · Statistics 2023-06-16 Jon A. Steingrimsson , Sarah E. Robertson , Issa J. Dahabreh

We consider a patient risk models which has access to patient features such as vital signs, lab values, and prior history but does not have access to a patient's diagnosis. For example, this occurs in a model deployed at intake time for…

Artificial Intelligence · Computer Science 2023-07-03 Alexander Peysakhovich , Rich Caruana , Yin Aphinyanaphongs