English
Related papers

Related papers: Selection and Collider Restriction Bias Due to Pre…

200 papers

A computable estimate of the readiness coefficient for a standard binary-state system is established in the case where both working and repair time distributions possess heavy tails.

Probability · Mathematics 2015-12-07 Alexander Veretennikov , Galina Zverkina

Unobserved confounding arises when an unmeasured feature influences both the treatment and the outcome, leading to biased causal effect estimates. This issue undermines observational studies in fields like economics, medicine, ecology or…

Machine Learning · Computer Science 2025-09-09 Alexander Merkov , David Rohde , Alexandre Gilotte , Benjamin Heymann

Outcome Reporting Bias (ORB) poses significant threats to the validity of meta-analytic findings. It occurs when researchers selectively report outcomes based on the significance or direction of results, potentially leading to distorted…

Methodology · Statistics 2025-07-17 Alessandra Gaia Saracini , Leonhard Held

We study randomized variants of two classical algorithms: coordinate descent for systems of linear equations and iterated projections for systems of linear inequalities. Expanding on a recent randomized iterated projection algorithm of…

Optimization and Control · Mathematics 2008-06-19 D. Leventhal , A. S. Lewis

The predictability of errors in deterministic temperature forecasts is investigated. More precisely, the aim is to issue warnings whenever the differences between forecast and verification exceed a given threshold. The warnings are…

Atmospheric and Oceanic Physics · Physics 2011-12-08 S. Hallerberg , J. Bröcker , H. Kantz , L. A. Smith

In the context of machine learning, disparate impact refers to a form of systematic discrimination whereby the output distribution of a model depends on the value of a sensitive attribute (e.g., race or gender). In this paper, we propose an…

Information Theory · Computer Science 2018-05-14 Hao Wang , Berk Ustun , Flavio P. Calmon

We propose a method for variable selection in discriminant analysis with mixed categorical and continuous variables. This method is based on a criterion that permits to reduce the variable selection problem to a problem of estimating…

Statistics Theory · Mathematics 2017-03-14 Alban Mbina Mbina , Guy Martial Nkiet , Fulgence Eyi Obiang

Models for predicting the time of a future event are crucial for risk assessment, across a diverse range of applications. Existing time-to-event (survival) models have focused primarily on preserving pairwise ordering of estimated event…

Machine Learning · Statistics 2021-01-14 Paidamoyo Chapfuwa , Chenyang Tao , Lawrence Carin , Ricardo Henao

The paper considers linear regression problems where the number of predictor variables is possibly larger than the sample size. The basic motivation of the study is to combine the points of view of model selection and functional regression…

Statistics Theory · Mathematics 2012-02-24 Alois Kneip , Pascal Sarda

Large differences between the properties of the known sample of cataclysmic variable stars (CVs) and the predictions of the theory of binary star evolution have long been recognised. However, because all existing CV samples suffer from…

Astrophysics · Physics 2008-11-26 Magaretha L. Pretorius , Christian Knigge , Ulrich Kolb

Classical epidemiology has focused on the control of confounding but it is only recently that epidemiologists have started to focus on the bias produced by colliders. A collider for a certain pair of variables (e.g., an outcome Y and an…

What we discover and see online, and consequently our opinions and decisions, are becoming increasingly affected by automated machine learned predictions. Similarly, the predictive accuracy of learning machines heavily depends on the…

Information Retrieval · Computer Science 2020-01-15 Sami Khenissi , Olfa Nasraoui

Two indicators are classically used to evaluate the quality of rule-based classification systems: predictive accuracy, i.e. the system's ability to successfully reproduce learning data and coverage, i.e. the proportion of possible cases for…

Artificial Intelligence · Computer Science 2020-04-07 Nassim Dehouche

In this note, we focus on a selection model problem: a mono-exponential model versus a bi-exponential one. This is done in the biological context of living cells, where small data are available. Classical statistics are revisited to improve…

Applications · Statistics 2011-05-31 Ph. Heinrich , J. Kahn , L. Héliot , D. Trinel

Covariate shift relaxes the widely-employed independent and identically distributed (IID) assumption by allowing different training and testing input distributions. Unfortunately, common methods for addressing covariate shift by trying to…

Machine Learning · Computer Science 2018-01-02 Anqi Liu , Brian D. Ziebart

The advent of powerful prediction algorithms led to increased automation of high-stake decisions regarding the allocation of scarce resources such as government spending and welfare support. This automation bears the risk of perpetuating…

Machine Learning · Statistics 2021-05-07 Matthias Kuppler , Christoph Kern , Ruben L. Bach , Frauke Kreuter

Scenario optimization and conformal prediction share a common goal, that is, turning finite samples into safety margins. Yet, different terminology often obscures the connection between their respective guarantees. This paper revisits that…

Systems and Control · Electrical Eng. & Systems 2026-03-23 Giuseppe C. Calafiore

Many popular algorithmic fairness measures depend on the joint distribution of predictions, outcomes, and a sensitive feature like race or gender. These measures are sensitive to distribution shift: a predictor which is trained to satisfy…

Machine Learning · Statistics 2022-02-11 Alan Mishler , Niccolò Dalmasso

Local sensitivity diagnostics for Bayesian models are described that are analogues of frequentist measures of leverage and influence. The diagnostics are simple to calculate using MCMC. A comparison between leverage and influence allows a…

Methodology · Statistics 2025-03-27 Martyn Plummer

Most data sets comprise of measurements on continuous and categorical variables. In regression and classification Statistics literature, modeling high-dimensional mixed predictors has received limited attention. In this paper we study the…

Statistics Theory · Mathematics 2021-10-26 Efstathia Bura , Liliana Forzani , Rodrigo García Arancibia , Pamela Llop , Diego Tomassi