English
Related papers

Related papers: The Illusion of Learning from Observational Data: …

200 papers

With the widespread accumulation of observational data, researchers obtain a new direction to learn counterfactual effects in many domains (e.g., health care and computational advertising) without Randomized Controlled Trials(RCTs).…

Machine Learning · Computer Science 2021-11-01 Guanglin Zhou , Lina Yao , Xiwei Xu , Chen Wang , Liming Zhu

The use of Bayesian information criterion (BIC) in the model selection procedure is under the assumption that the observations are independent and identically distributed (i.i.d.). However, in practice, we do not always have i.i.d. samples.…

Applications · Statistics 2021-05-03 Nan Shen , Bárbara González

We consider the problem of online learning in the presence of distribution shifts that occur at an unknown rate and of unknown intensity. We derive a new Bayesian online inference approach to simultaneously infer these distribution shifts…

Machine Learning · Statistics 2021-10-28 Aodong Li , Alex Boyd , Padhraic Smyth , Stephan Mandt

Observed associations in a database may be due in whole or part to variations in unrecorded (latent) variables. Identifying such variables and their causal relationships with one another is a principal goal in many scientific and practical…

Machine Learning · Computer Science 2012-12-12 Ricardo Silva , Richard Scheines , Clark Glymour , Peter L. Spirtes

While many areas of machine learning have benefited from the increasing availability of large and varied datasets, the benefit to causal inference has been limited given the strong assumptions needed to ensure identifiability of causal…

Machine Learning · Computer Science 2022-01-02 Wenshuo Guo , Serena Wang , Peng Ding , Yixin Wang , Michael I. Jordan

Causal learning is the cognitive process of developing the capability of making causal inferences based on available information, often guided by normative principles. This process is prone to errors and biases, such as the illusion of…

We consider causal models with two observed variables and one latent variables, each variable being discrete, with the goal of characterizing the possible distributions on outcomes that can result from controlling one of the observed…

Information Theory · Computer Science 2021-03-05 Kevin Shu

Gaussian empirical Bayes methods usually maintain a precision independence assumption: The unknown parameters of interest are independent from the known standard errors of the estimates. This assumption is often theoretically questionable…

Econometrics · Economics 2025-12-30 Jiafeng Chen

This study introduces a data-driven, machine learning-based method to detect suitable control variables and instruments for assessing the causal effect of a treatment on an outcome in observational data. Our approach tests the joint…

Econometrics · Economics 2026-05-20 Nicolas Apfel , Julia Hatamyar , Martin Huber , Jannis Kueck

We derive a Bayesian framework for incorporating selection effects into population analyses. We allow for both measurement uncertainty in individual measurements and, crucially, for selection biases on the population of measurements, and…

Data Analysis, Statistics and Probability · Physics 2019-04-10 Ilya Mandel , Will M. Farr , Jonathan R. Gair

We study the convergence rates of empirical Bayes posterior distributions for nonparametric and high-dimensional inference. We show that as long as the hyperparameter set is discrete, the empirical Bayes posterior distribution induced by…

Statistics Theory · Mathematics 2020-09-10 Fengshuo Zhang , Chao Gao

When an analyst or scientist has a belief about how the world works, their thinking can be biased in favor of that belief. Therefore, one bedrock principle of science is to minimize that bias by testing the predictions of one's belief…

Human-Computer Interaction · Computer Science 2022-08-10 Cindy Xiong , Chase Stokes , Yea-Seul Kim , Steven Franconeri

One obstacle to ``elevating" correlation to causation is the phenomenon of confounding, i.e., when a correlation between two variables exists because both variables are in fact caused by a third variable. The situation where the confounders…

Applications · Statistics 2025-06-24 Caren Marzban , Yikun Zhang , Nicholas Bond , Michael Richman

There is intense interest in applying machine learning to problems of causal inference in fields such as healthcare, economics and education. In particular, individual-level causal inference has important applications such as precision…

Machine Learning · Statistics 2017-05-17 Uri Shalit , Fredrik D. Johansson , David Sontag

Recent work introduced loss functions which measure the error of a prediction based on multiple simultaneous observations or outcomes. In this paper, we explore the theoretical and practical questions that arise when using such…

Machine Learning · Computer Science 2018-02-28 Rafael Frongillo , Nishant A. Mehta , Tom Morgan , Bo Waggoner

Learning representations purely from observations concerns the problem of learning a low-dimensional, compact representation which is beneficial to prediction models. Under the hypothesis that the intrinsic latent factors follow some casual…

Machine Learning · Computer Science 2023-10-24 Mengyue Yang , Xinyu Cai , Furui Liu , Weinan Zhang , Jun Wang

Bayesian inference gets its name from *Bayes's theorem*, expressing posterior probabilities for hypotheses about a data generating process as the (normalized) product of prior probabilities and a likelihood function. But Bayesian inference…

Methodology · Statistics 2024-07-02 Thomas J. Loredo , Robert L. Wolpert

Causal inference in observational studies is notoriously difficult, due to the fact that the experimenter is not in charge of the treatment assignment mechanism. Many potential con- founding factors (PCFs) exist in such a scenario, and if…

Applications · Statistics 2015-09-15 Vadim von Brzeski , Matt Taddy , David Draper

A new empirical Bayes approach to variable selection in the context of generalized linear models is developed. The proposed algorithm scales to situations in which the number of putative explanatory variables is very large, possibly much…

Methodology · Statistics 2021-06-29 Haim Bar , James Booth , Martin T. Wells

Most positive and unlabeled data is subject to selection biases. The labeled examples can, for example, be selected from the positive set because they are easier to obtain or more obviously positive. This paper investigates how learning can…

Machine Learning · Computer Science 2019-07-01 Jessa Bekker , Pieter Robberechts , Jesse Davis