English
Related papers

Related papers: Contextualizing selection bias in Mendelian random…

200 papers

This paper addresses the sample selection model within the context of the gender gap problem, where even random treatment assignment is affected by selection bias. By offering a robust alternative free from distributional or specification…

Econometrics · Economics 2024-10-04 Xiaolin Sun , Xueyan Zhao , D. S. Poskitt

Bayesian model comparison is often based on the posterior distribution over the set of compared models. This distribution is often observed to concentrate on a single model even when other measures of model fit or forecasting ability…

Statistics Theory · Mathematics 2020-03-10 Oscar Oelrich , Shutong Ding , Måns Magnusson , Aki Vehtari , Mattias Villani

Phylodynamics seeks to estimate effective population size fluctuations from molecular sequences of individuals sampled from a population of interest. One way to accomplish this task formulates an observed sequence data likelihood exploiting…

We consider large-scale studies in which it is of interest to test a very large number of hypotheses, and then to estimate the effect sizes corresponding to the rejected hypotheses. For instance, this setting arises in the analysis of gene…

Methodology · Statistics 2015-03-31 Kean Ming Tan , Noah Simon , Daniela Witten

"M-Bias," as it is called in the epidemiologic literature, is the bias introduced by conditioning on a pretreatment covariate due to a particular "M-Structure" between two latent factors, an observed treatment, an outcome, and a "collider."…

Statistics Theory · Mathematics 2014-08-05 Peng Ding , Luke Miratrix

The literature on cluster-randomized trials typically allows for interference within but not across clusters. This may be implausible when units are irregularly distributed across space without well-separated communities, as clusters in…

Methodology · Statistics 2025-10-29 Michael P. Leung

Bias can be introduced in diverse ways in machine learning datasets, for example via selection or label bias. Although these bias types in themselves have an influence on important aspects of fair machine learning, their different impact…

Machine Learning · Computer Science 2026-03-11 Magali Legast , Toon Calders , François Fouss

Prediction with the possibility of abstention (or selective prediction) is an important problem for error-critical machine learning applications. While well-studied in the classification setup, selective approaches to regression are much…

Machine Learning · Statistics 2023-09-29 Fedor Noskov , Alexander Fishkov , Maxim Panov

In the causal adjustment setting, variable selection techniques based on either the outcome or treatment allocation model can result in the omission of confounders or the inclusion of spurious variables in the propensity score. We propose a…

Statistics Theory · Mathematics 2014-06-06 Ashkan Ertefaie , Masoud Asgharian , David A. Stephens

Mendelian randomization (MR) is a statistical method exploiting genetic variants as instrumental variables to estimate the causal effect of modifiable risk factors on an outcome of interest. Despite wide uses of various popular two-sample…

Methodology · Statistics 2021-11-17 Anqi Wang , Zhonghua Liu

Results in epidemiology and social science often require the removal of confounding effects from measurements of the pairwise correlation of variables in survey data. This is typically accomplished by some variant of linear regression…

Methodology · Statistics 2025-12-02 William H. Press

A fundamental component in the theoretical school choice literature is the problem a student faces in deciding which schools to apply to. Recent models have considered a set of schools of different selectiveness and a student who is unsure…

Computer Science and Game Theory · Computer Science 2024-03-08 Jon Kleinberg , Sigal Oren , Emily Ryu , Éva Tardos

Observational studies are needed when experiments are not possible. Within study comparisons (WSC) compare observational and experimental estimates that test the same hypothesis using the same treatment group, outcome, and estimand.…

Medical studies for chronic disease are often interested in the relation between longitudinal risk factor profiles and individuals' later life disease outcomes. These profiles may typically be subject to intermediate structural changes due…

Applications · Statistics 2023-08-22 Sandra Keizer , Zhuozhao Zhan , Vasan S. Ramachandran , Edwin R. van den Heuvel

Clustering is a central approach for unsupervised learning. After clustering is applied, the most fundamental analysis is to quantitatively compare clusterings. Such comparisons are crucial for the evaluation of clustering methods as well…

Machine Learning · Statistics 2017-10-03 Alexander J Gates , Yong-Yeol Ahn

Mendelian diseases are determined by a single mutation in a given gene. However, in the case of diseases with late onset, the age at onset is variable; it can even be the case that the onset is not observed in a lifetime. Estimating the…

Applications · Statistics 2016-07-15 Flora Alarcon , Gregory Nuel , Violaine Plante-Bordeneuve

Multivariate meta-analysis (MMA) is a powerful tool for jointly estimating multiple outcomes' treatment effects. However, the validity of results from MMA is potentially compromised by outcome reporting bias (ORB), or the tendency for…

Applications · Statistics 2021-10-19 Ray Bai , Xiaokang Liu , Lifeng Lin , Yulun Liu , Stephen E. Kimmel , Haitao Chu , Yong Chen

Variable selection, or more generally, model reduction is an important aspect of the statistical workflow aiming to provide insights from data. In this paper, we discuss and demonstrate the benefits of using a reference model in variable…

Methodology · Statistics 2020-04-29 Federico Pavone , Juho Piironen , Paul-Christian Bürkner , Aki Vehtari

We discuss the role of misspecification and censoring on Bayesian model selection in the contexts of right-censored survival and concave log-likelihood regression. Misspecification includes wrongly assuming the censoring mechanism to be…

Methodology · Statistics 2021-11-15 David Rossell , Francisco Javier Rubio

The purpose of modeling document relevance for search engines is to rank better in subsequent searches. Document-specific historical click-through rates can be important features in a dynamic ranking system which updates as we accumulate…

Information Retrieval · Computer Science 2024-02-06 Richard Demsyn-Jones