English
Related papers

Related papers: Prediction-powered Generalization of Causal Infere…

200 papers

This study aims to understand how statistical biases affect the model's ability to generalize to in-distribution and out-of-distribution data on algorithmic tasks. Prior research indicates that transformers may inadvertently learn to rely…

Machine Learning · Computer Science 2024-09-11 John Mitros

Randomized controlled trials often enroll participants whose characteristics differ from those of a target population, which can limit the generalizability of the estimated treatment effects when effect modifiers differ across populations.…

Methodology · Statistics 2026-05-15 Lan Wen , Issa J. Dahabreh , Yu-Han Chiu

Reinforcement learning enhances the reasoning capabilities of large language models but often involves high computational costs due to rollout-intensive optimization. Online prompt selection presents a plausible solution by prioritizing…

Artificial Intelligence · Computer Science 2026-05-18 Yun Qu , Qi Wang , Yixiu Mao , Heming Zou , Yuhang Jiang , Weijie Liu , Clive Bai , Kai Yang , Yangkun Chen , Saiyong Yang , Xiangyang Ji

Although randomized controlled trials (RCTs) are a cornerstone of comparative effectiveness, they typically have much smaller sample size than observational studies because of financial and ethical considerations. Therefore there is…

Methodology · Statistics 2023-11-16 Lauren D. Liao , Emilie Højbjerre-Frandsen , Alan E. Hubbard , Alejandro Schuler

Randomized clinical trials are considered the gold standard for informing treatment guidelines, but results may not generalize to real-world populations. Generalizability is hindered by distributional differences in baseline covariates and…

Methodology · Statistics 2025-06-03 Rachael K. Ross , Ivan Diaz , Amy J. Pitts , Elizabeth A. Stuart , Kara E. Rudolph

Clinical trials often collect data on multiple outcomes, such as overall survival (OS), progression-free survival (PFS), and response to treatment (RT). In most cases, however, study designs only use primary outcome data for interim and…

Applications · Statistics 2026-04-28 Massimiliano Russo , Steffen Ventz , Lorenzo Trippa

Real-World Data (RWD), with its large sample sizes and rich clinical detail, offers a compelling alternative to randomized controlled trials (RCTs) for studying treatment effects in diverse and complex patient populations. However, its…

Applications · Statistics 2026-05-26 Yifei Xu , Hwiyoung Lee , Zhenyao Ye , Yezhi Pan , Jingsong Zhou , Yun Yang , Chixiang Chen , Shuo Chen

Objective: Randomized controlled trial (RCT) results often inform clinical decision-making, but the highly curated populations of trials and the care provided during the trial are often not reflective of real-world practice. The objective…

Applications · Statistics 2024-02-26 Guanbo Wang , Ting-Wei Ernie Liao , David Furfaro , Leo Anthony Celi , Kevin Sheng-Kai Ma

The inaccessibility of controlled randomized trials due to inherent constraints in many fields of science has been a fundamental issue in causal inference. In this paper, we focus on distinguishing the cause from effect in the bivariate…

Machine Learning · Statistics 2021-02-23 Jean-Francois Ton , Dino Sejdinovic , Kenji Fukumizu

Prediction and causal explanation are fundamentally distinct tasks of data analysis. In health applications, this difference can be understood in terms of the difference between prognosis (prediction) and prevention/treatment (causal…

In the realm of stock prediction, machine learning models encounter considerable obstacles due to the inherent low signal-to-noise ratio and the nonstationary nature of financial markets. These challenges often result in spurious…

Portfolio Management · Quantitative Finance 2025-03-28 Songci Xu , Qiangqiang Cheng , Chi-Guhn Lee

Machine learning applications often require calibrated predictions, e.g. a 90\% credible interval should contain the true outcome 90\% of the times. However, typical definitions of calibration only require this to hold on average, and offer…

Machine Learning · Statistics 2020-09-10 Shengjia Zhao , Tengyu Ma , Stefano Ermon

The purpose of many health studies is to estimate the effect of an exposure on an outcome. It is not always ethical to assign an exposure to individuals in randomised controlled trials, instead observational data and appropriate study…

One of the principal scientific challenges in deep learning is explaining generalization, i.e., why the particular way the community now trains networks to achieve small training error also leads to small error on held-out data from the…

Recent methods to improve generalizations from nonrandom samples typically invoke assumptions such as the strong ignorability of sample selection that are often controversial in practice to derive point estimates. Rather than focus on the…

Applications · Statistics 2017-01-06 Wendy Chan

Randomized trials typically estimate average relative treatment effects, but decisions on the benefit of a treatment are possibly better informed by more individualized predictions of the absolute treatment effect. In case of a binary…

Methodology · Statistics 2021-08-20 J Hoogland , J IntHout , M Belias , MM Rovers , RD Riley , FE Harrell , KGM Moons , TPA Debray , JB Reitsma

Prior work has shown that language models can be tuned to follow user instructions using only a small set of high-quality instructions. This has accelerated the development of methods that filter a large, noisy instruction-tuning datasets…

Artificial Intelligence · Computer Science 2024-10-22 Harshita Diddee , Daphne Ippolito

Deep neural networks have achieved remarkable success in various challenging tasks. However, the black-box nature of such networks is not acceptable to critical applications, such as healthcare. In particular, the existence of adversarial…

Machine Learning · Computer Science 2019-09-12 Shaeke Salman , Seyedeh Neelufar Payrovnaziri , Xiuwen Liu , Pablo Rengifo-Moreno , Zhe He

A predictive model makes outcome predictions based on some given features, i.e., it estimates the conditional probability of the outcome given a feature vector. In general, a predictive model cannot estimate the causal effect of a feature…

Machine Learning · Computer Science 2023-04-11 Jiuyong Li , Lin Liu , Ziqi Xu , Ha Xuan Tran , Thuc Duy Le , Jixue Liu

Learned systems in the domain of visual recognition and cognition impress in part because even though they are trained with datasets many orders of magnitude smaller than the full population of possible images, they exhibit sufficient…

Computer Vision and Pattern Recognition · Computer Science 2022-05-03 John K. Tsotsos , Jun Luo