English
Related papers

Related papers: SelectionBias: An R Package for Bounding Selection…

200 papers

Due to its linear complexity, naive Bayes classification remains an attractive supervised learning method, especially in very large-scale settings. We propose a sparse version of naive Bayes, which can be used for feature selection. This…

Machine Learning · Computer Science 2025-03-13 Armin Askari , Alexandre d'Aspremont , Laurent El Ghaoui

Selection bias is pervasive in observational studies. For example, large scale biobanks data can exhibit ``healthy volunteer bias'' when respondents are healthier and of higher socio-economic status than the population they are meant to…

Methodology · Statistics 2026-05-14 Yiwen Qiu , Filip Kovacevic , Shimeng Huang , Peter Spirtes , Francesco Locatello

Model selection aims to identify a sufficiently well performing model that is possibly simpler than the most complex model among a pool of candidates. However, the decision-making process itself can inadvertently introduce non-negligible…

Methodology · Statistics 2024-08-08 Yann McLatchie , Aki Vehtari

Sentiment Analysis (SA) refers to a family of techniques at the crossroads of statistics, natural language processing, and computational linguistics. The primary goal is to detect the semantic orientation of individual opinions and comments…

Information Retrieval · Computer Science 2026-03-09 Michelangelo Misuraca , Alessia Forciniti , Germana Scepi , Maria Spano

In observational studies of treatment effects, estimates may be biased by unmeasured confounders, which can potentially affect the validity of the results. Understanding sensitivity to such biases helps assess how unmeasured confounding…

Methodology · Statistics 2026-02-02 Qishuo Yin , Dylan S. Small

Estimating the causal effect of a treatment or health policy with observational data can be challenging due to an imbalance of and a lack of overlap between treated and control covariate distributions. In the presence of limited overlap,…

Methodology · Statistics 2025-03-24 Martha Barnard , Jared D. Huling , Julian Wolfson

We study three classical machine learning algorithms in the context of algorithmic fairness: adaptive boosting, support vector machines, and logistic regression. Our goal is to maintain the high accuracy of these learning algorithms while…

Machine Learning · Computer Science 2016-01-22 Benjamin Fish , Jeremy Kun , Ádám D. Lelkes

Many measurements at collider experiments study physics candidates that are a subset of a collision event. The presence of multiple such candidates in a given event can cause raw biases which are large compared to typical statistical…

High Energy Physics - Experiment · Physics 2019-08-22 Patrick Koppenburg

We consider the problem of estimating the expected causal effect $E[Y|do(X)]$ for a target variable $Y$ when treatment $X$ is set by intervention, focusing on continuous random variables. In settings without selection bias or confounding,…

Machine Learning · Statistics 2025-03-27 Marlies Hafer , Alexander Marx

The statistics and machine learning communities have recently seen a growing interest in classification-based approaches to two-sample testing. The outcome of a classification-based two-sample test remains a rejection decision, which is not…

Statistics Theory · Mathematics 2022-11-15 Loris Michel , Jeffrey Näf , Nicolai Meinshausen

The choice of the prior distribution is a key aspect of Bayesian analysis. For the spatial regression setting a subjective prior choice for the parameters may not be trivial, from this perspective, using the objective Bayesian analysis…

Statistics Theory · Mathematics 2020-04-10 Jose A. Ordoñez , Marcos O. Prates , Larissa A. Matos , Victor H. Lachos

To estimate direct and indirect effects of an exposure on an outcome from observed data strong assumptions about unconfoundedness are required. Since these assumptions cannot be tested using the observed data, a mediation analysis should…

Statistics Theory · Mathematics 2018-03-29 Anita Lindmark , Xavier de Luna , Marie Eriksson

We focus on a stochastic learning model where the learner observes a finite set of training examples and the output of the learning process is a data-dependent distribution over a space of hypotheses. The learned data-dependent distribution…

Machine Learning · Statistics 2020-12-29 Omar Rivasplata , Ilja Kuzborskij , Csaba Szepesvari , John Shawe-Taylor

Contemporary sample size calculations for external validation of risk prediction models require users to specify fixed values of assumed model performance metrics alongside target precision levels (e.g., 95% CI widths). However, due to the…

Applications · Statistics 2026-02-13 Mohsen Sadatsafavi , Paul Gustafson , Solmaz Setayeshgar , Laure Wynants , Richard D Riley

Sensitivity Analysis is a framework to assess how conclusions drawn from missing outcome data may be vulnerable to departures from untestable underlying assumptions. We extend the E-value, a popular metric for quantifying robustness of…

Methodology · Statistics 2021-08-31 Wu Xue , Abbas Zaidi

With the aim of building machine learning systems that incorporate standards of fairness and accountability, we explore explicit subgroup sample complexity bounds. The work is motivated by the observation that classifier predictions for…

Machine Learning · Computer Science 2019-10-28 Ananth Balashankar , Alyssa Lees

Researchers would often like to leverage data from a collection of sources (e.g., primary studies in a meta-analysis) to estimate causal effects in a target population of interest. However, traditional meta-analytic methods do not produce…

Methodology · Statistics 2025-05-15 Guanbo Wang , Sean McGrath , Yi Lian

Instrumental variables regression is a tool that is commonly used in the analysis of observational data. The instrumental variables are used to make causal inference about the effect of a certain exposure in the presence of unmeasured…

Methodology · Statistics 2023-09-07 Valentin Vancak , Arvid Sjölander

While robust divergence such as density power divergence and $\gamma$-divergence is helpful for robust statistical inference in the presence of outliers, the tuning parameter that controls the degree of robustness is chosen in a…

Methodology · Statistics 2021-09-15 Shonosuke Sugasawa , Shouto Yonekura

We present a novel analysis of the expected risk of weighted majority vote in multiclass classification. The analysis takes correlation of predictions by ensemble members into account and provides a bound that is amenable to efficient…

Machine Learning · Computer Science 2020-12-18 Andrés R. Masegosa , Stephan S. Lorenzen , Christian Igel , Yevgeny Seldin
‹ Prev 1 4 5 6 7 8 10 Next ›