中文
相关论文

相关论文: Contextualizing selection bias in Mendelian random…

200 篇论文

External controls from historical trials or observational data can augment randomized controlled trials when large-scale randomization is impractical or unethical, such as in drug evaluation for rare diseases. However, non-randomized…

统计方法学 · 统计学 2025-05-08 Ke Zhu , Shu Yang , Xiaofei Wang

We propose a novel definition of selection bias in analytic epidemiology using potential outcomes. This definition captures selection bias under both the structural approach (where conditioning on selection into the study opens a noncausal…

统计方法学 · 统计学 2023-10-24 Eben Kenah

When using machine learning for imbalanced binary classification problems, it is common to subsample the majority class to create a (more) balanced training dataset. This biases the model's predictions because the model learns from data…

机器学习 · 计算机科学 2025-11-03 Nathan Phelps , Daniel J. Lizotte , Douglas G. Woolford

In observational studies, researchers must select a method to control for confounding. Options include propensity score methods and regression. It remains unclear how dataset characteristics (size, overlap in propensity scores, exposure…

统计方法学 · 统计学 2022-10-21 J. Wilkinson , M. A. Mamas , E. Kontopantelis

Estimating the causal effect of a treatment or health policy with observational data can be challenging due to an imbalance of and a lack of overlap between treated and control covariate distributions. In the presence of limited overlap,…

统计方法学 · 统计学 2025-03-24 Martha Barnard , Jared D. Huling , Julian Wolfson

We derive a family of loss functions to train models in the presence of sampling bias. Examples are when the prevalence of a pathology differs from its sampling rate in the training dataset, or when a machine learning practioner rebalances…

Prediction for very large data sets is typically carried out in two stages, variable selection and pattern recognition. Ordinarily variable selection involves seeing how well individual explanatory variables are correlated with the…

统计方法学 · 统计学 2017-09-12 Herman Chernoff , Shaw-Hwa Lo , Tian Zheng , Adeline Lo

Large Language Models (LLMs) are widely used to evaluate natural language generation tasks as automated metrics. However, the likelihood, a measure of LLM's plausibility for a sentence, can vary due to superficial differences in sentences,…

计算与语言 · 计算机科学 2025-11-11 Masanari Oi , Masahiro Kaneko , Ryuto Koike , Mengsay Loem , Naoaki Okazaki

Inferring the causal effect of a treatment on an outcome in an observational study requires adjusting for observed baseline confounders to avoid bias. However, adjusting for all observed baseline covariates, when only a subset are…

统计方法学 · 统计学 2021-02-04 Wen Wei Loh , Stijn Vansteelandt

Many diseases and traits involve a complex interplay between genes and environment, generating significant interest in studying gene-environment interaction through observational data. However, for lifestyle and environmental risk factors,…

统计方法学 · 统计学 2023-09-22 Malka Gorfine , Conghui Qu , Ulrike Peters , Li Hsu

We propose a cautious Bayesian variable selection routine by investigating the sensitivity of a hierarchical model, where the regression coefficients are specified by spike and slab priors. We exploit the use of latent variables to…

统计方法学 · 统计学 2022-06-20 Tathagata Basu , Matthias C. M. Troffaes , Jochen Einbeck

The Bayesian method is noted to produce spuriously high posterior probabilities for phylogenetic trees in analysis of large datasets, but the precise reasons for this over-confidence are unknown. In general, the performance of Bayesian…

统计理论 · 数学 2018-10-15 Ziheng Yang , Tianqi Zhu

Background: The hazard ratio of the Cox proportional hazards model is widely used in randomized controlled trials to assess treatment effects. However, two properties of the hazard ratio including the non-collapsibility and built-in…

统计方法学 · 统计学 2024-01-17 Helen Bian , Menglan Pang , Guanbo Wang , Zihang Lu

Standard methods in preference learning involve estimating the parameters of discrete choice models from data of selections (choices) made by individuals from a discrete set of alternatives (the choice set). While there are many models for…

机器学习 · 计算机科学 2021-08-18 Kiran Tomlinson , Johan Ugander , Austin R. Benson

We show that if the exposure and the outcome affect the selection indicator in the same direction and have non-positive interaction on the risk difference, risk ratio or odds ratio scale, the exposure-outcome odds ratio in the selected…

统计理论 · 数学 2017-01-17 Zhichao Jiang , Peng Ding

We study a population of $N$ individuals evolving according to a biparental Moran model with two types, one being advantaged compared to the other. The advantage is conferred by a Mendelian mutation, which reduces the death probability of…

概率论 · 数学 2026-03-24 Camille Coron , Yves Le Jan

Machine learning models are susceptible to being misled by biases in training data that emphasize incidental correlations over the intended learning task. In this study, we demonstrate the impact of data bias on the performance of a machine…

材料科学 · 物理学 2024-12-11 Ali Davariashtiyani , Busheng Wang , Samad Hajinazar , Eva Zurek , Sara Kadkhodaei

We address the common yet often-overlooked selection bias in interventional studies, where subjects are selectively enrolled into experiments. For instance, participants in a drug trial are usually patients of the relevant disease; A/B…

机器学习 · 计算机科学 2025-03-11 Haoyue Dai , Ignavier Ng , Jianle Sun , Zeyu Tang , Gongxu Luo , Xinshuai Dong , Peter Spirtes , Kun Zhang

Mendelian Randomisation (MR) uses genetic variants as instrumental variables to infer causal effects of exposures on an outcome. One key assumption of MR is that the genetic variants used as instrumental variables are independent of the…

统计方法学 · 统计学 2025-02-21 Maximilian M Mandl , Anne-Laure Boulesteix , Stephen Burgess , Verena Zuber

Background: Mendelian randomization (MR) has been widely applied to causal inference in medical research. It uses genetic variants as instrumental variables (IVs) to investigate putative causal relationship between an exposure and an…

统计方法学 · 统计学 2020-11-04 Linyi Zou , Hui Guo , Carlo Berzuini