English
Related papers

Related papers: Robustifying and Selecting Cohort-Appropriate Prog…

200 papers

Generative models frequently suffer miscalibration, wherein statistics of the sampling distribution, such as the fraction of generations in a given class, deviate from desired values. We frame calibration as a constrained optimization…

Machine Learning · Statistics 2026-05-29 Henry D. Smith , Nathaniel L. Diamant , Brian L. Trippe

In observational causal inference, domain knowledge often leaves multiple covariate adjustments plausible, yet which sets satisfy ignorability is untestable. Different adjustment sets can yield conflicting estimates of the average treatment…

Methodology · Statistics 2026-03-23 Aditya Ghosh , Dominik Rothenhäusler

Suppose we are interested in the effect of a treatment in a clinical trial. The efficiency of inference may be limited due to small sample size. However, external control data are often available from historical studies. Motivated by an…

Methodology · Statistics 2021-12-10 Xinyu Li , Wang Miao , Fang Lu , Xiao-Hua Zhou

In cluster-randomized trials, generalized linear mixed models and generalized estimating equations have conventionally been the default analytic methods for estimating the average treatment effect as routine practice. However, recent…

Methodology · Statistics 2025-09-19 Fan Li , Jiaqi Tong , Xi Fang , Chao Cheng , Brennan C. Kahan , Bingkai Wang

In federated learning, differences in the data or objectives between the participating nodes motivate approaches to train a personalized machine learning model for each node. One such approach is weighted averaging between a locally trained…

Machine Learning · Computer Science 2021-10-26 Felix Grimberg , Mary-Anne Hartley , Sai P. Karimireddy , Martin Jaggi

In recent years, real-world external controls have grown in popularity as a tool to empower randomized placebo-controlled trials, particularly in rare diseases or cases where balanced randomization is unethical or impractical. However, as…

Methodology · Statistics 2024-11-14 Chenyin Gao , Shu Yang , Mingyang Shan , Wenyu Ye , Ilya Lipkovich , Douglas Faries

Covariate adjustment is a general method for improving precision when estimating treatment effects in randomized trials and is recommended by the FDA in its 2023 guidance when baseline variables are prognostic for the primary outcome. We…

We are concerned with obtaining well-calibrated output distributions from regression models. Such distributions allow us to quantify the uncertainty that the model has regarding the predicted target value. We introduce the novel concept of…

Machine Learning · Statistics 2019-05-16 Hao Song , Tom Diethe , Meelis Kull , Peter Flach

There is a growing need for flexible general frameworks that integrate individual-level data with external summary information for improved statistical inference. External information relevant for a risk prediction model may come in…

Methodology · Statistics 2023-04-11 Tian Gu , Jeremy M. G. Taylor , Bhramar Mukherjee

A prediction model is most useful if it generalizes beyond the development data with external validations, but to what extent should it generalize remains unclear. In practice, prediction models are externally validated using data from very…

Machine Learning · Computer Science 2023-04-11 Yilin Ning , Victor Volovici , Marcus Eng Hock Ong , Benjamin Alan Goldstein , Nan Liu

Machine learning models traditionally assume that training and test data are independently and identically distributed. However, in real-world applications, the test distribution often differs from training. This problem, known as…

Machine Learning · Computer Science 2024-06-19 Kotaro Yoshida , Hiroki Naganuma

Machine learning applications often require calibrated predictions, e.g. a 90\% credible interval should contain the true outcome 90\% of the times. However, typical definitions of calibration only require this to hold on average, and offer…

Machine Learning · Statistics 2020-09-10 Shengjia Zhao , Tengyu Ma , Stefano Ermon

Randomized trials are widely considered as the gold standard for evaluating the effects of decision policies. Trial data is, however, drawn from a population which may differ from the intended target population and this raises a problem of…

Methodology · Statistics 2024-10-30 Sofia Ek , Dave Zachariah

Estimating the Kullback-Leibler (KL) divergence between two distributions given samples from them is well-studied in machine learning and information theory. Motivated by considerations of multi-group fairness, we seek KL divergence…

Machine Learning · Computer Science 2022-03-01 Parikshit Gopalan , Nina Narodytska , Omer Reingold , Vatsal Sharan , Udi Wieder

Usually in Latent Class Analysis (LCA), external predictors are taken to be cluster conditional probability predictors (LC models with covariates), and/or score conditional probability predictors (LC regression models). In such cases, their…

Methodology · Statistics 2018-01-08 Roberto Di Mari , Antonio Punzo , Zsuzsa Bakk

With model trustworthiness being crucial for sensitive real-world applications, practitioners are putting more and more focus on improving the uncertainty calibration of deep neural networks. Calibration errors are designed to quantify the…

Machine Learning · Computer Science 2024-03-14 Sebastian G. Gruber , Florian Buettner

The partitioning of data for estimation and calibration critically impacts the performance of propensity score based estimators like inverse probability weighting (IPW) and double/debiased machine learning (DML) frameworks. We extend recent…

Machine Learning · Statistics 2025-05-20 Sven Klaassen , Jan Rabenseifner , Jannis Kueck , Philipp Bach

An accurate model of a patient's individual survival distribution can help determine the appropriate treatment for terminal patients. Unfortunately, risk scores (e.g., from Cox Proportional Hazard models) do not provide survival…

Machine Learning · Computer Science 2020-07-08 Humza Haider , Bret Hoehn , Sarah Davis , Russell Greiner

Randomized clinical trials are considered the gold standard for informing treatment guidelines, but results may not generalize to real-world populations. Generalizability is hindered by distributional differences in baseline covariates and…

Methodology · Statistics 2025-06-03 Rachael K. Ross , Ivan Diaz , Amy J. Pitts , Elizabeth A. Stuart , Kara E. Rudolph

Continual Learning (CL) focuses on maximizing the predictive performance of a model across a non-stationary stream of data. Unfortunately, CL models tend to forget previous knowledge, thus often underperforming when compared with an offline…

Machine Learning · Computer Science 2024-04-15 Lanpei Li , Elia Piccoli , Andrea Cossu , Davide Bacciu , Vincenzo Lomonaco