English
Related papers

Related papers: The effect of collinearity and sample size on line…

200 papers

Unlike parametric regression, machine learning (ML) methods do not generally require precise knowledge of the true data generating mechanisms. As such, numerous authors have advocated for ML methods to estimate causal effects.…

Methodology · Statistics 2020-05-15 Ashley I Naimi , Alan E Mishler , Edward H Kennedy

We present a randomized controlled trial for a model-in-the-loop regression task, with the goal of measuring the extent to which (1) good explanations of model predictions increase human accuracy, and (2) faulty explanations decrease human…

Machine Learning · Computer Science 2020-07-27 Eric Chu , Deb Roy , Jacob Andreas

In practical regression applications, multiple covariates are often measured, but not all may be associated with the response variable. Identifying and including only the relevant covariates in the model is crucial for improving prediction…

Methodology · Statistics 2026-03-10 Ana Carolina da Cruz , Camila P. E. de Souza , Pedro H. T. O. Sousa

We consider the identification of large-scale linear and stable dynamic systems whose outputs may be the result of many correlated inputs. Hence, severe ill-conditioning may affect the estimation problem. This is a scenario often arising…

Dynamical Systems · Mathematics 2022-09-07 Wenqi Cao , Gianluigi Pillonetto

Linear mixed-effects models are widely used in analyzing repeated measures data, including clustered and longitudinal data, where inferences of both fixed effects and variance components are of importance. Unlike the fixed effect inference…

Methodology · Statistics 2022-08-31 J. Zhang , W. Guo , J. S. Carpenter , Andrew Leroux , K. R. Merikangas , N. G. Martin , I. B. Hickie , H. Shou , H. Li

As a widely used weakly supervised learning scheme, modern multiple instance learning (MIL) models achieve competitive performance at the bag level. However, instance-level prediction, which is essential for many important applications,…

Machine Learning · Computer Science 2022-06-14 Hitesh Sapkota , Qi Yu

The multivariable fractional polynomial (MFP) procedure combines variable selection with a function selection procedure (FSP). For continuous variables, a closed test procedure is used to decide between no effect, linear, FP1 or FP2…

Methodology · Statistics 2022-09-21 Willi Sauerbrei , Edwin Kipruto , James Balmford

Causal effect estimation from observational data is an important and much studied research topic. The instrumental variable (IV) and local causal discovery (LCD) patterns are canonical examples of settings where a closed-form expression…

Machine Learning · Statistics 2018-09-19 Ioan Gabriel Bucur , Tom Claassen , Tom Heskes

Instrumental variables (IV) estimation suffers selection bias when the analysis conditions on the treatment. Judea Pearl's early graphical definition of instrumental variables explicitly prohibited conditioning on the treatment.…

Econometrics · Economics 2020-05-20 Felix Elwert , Elan Segarra

One of the core problems of modern statistics is to approximate difficult-to-compute probability densities. This problem is especially important in Bayesian statistics, which frames all inference about unknown quantities as a calculation…

Computation · Statistics 2018-05-11 David M. Blei , Alp Kucukelbir , Jon D. McAuliffe

When using machine learning for imbalanced binary classification problems, it is common to subsample the majority class to create a (more) balanced training dataset. This biases the model's predictions because the model learns from data…

Machine Learning · Computer Science 2025-11-03 Nathan Phelps , Daniel J. Lizotte , Douglas G. Woolford

Conformal prediction, which makes no distributional assumptions about the data, has emerged as a powerful and reliable approach to uncertainty quantification in practical applications. The nonconformity measure used in conformal prediction…

Machine Learning · Computer Science 2024-10-15 Yuko Kato , David M. J. Tax , Marco Loog

Model checking plays an important role in linear regression as model misspecification seriously affects the validity and efficiency of regression analysis. In practice, model checking is often performed by subjectively evaluating the plot…

Statistics Theory · Mathematics 2019-11-19 Rok Blagus , Jakob Peterlin , Janez Stare

Class imbalanced datasets are common in real-world applications that range from credit card fraud detection to rare disease diagnostics. Several popular classification algorithms assume that classes are approximately balanced, and hence…

Machine Learning · Statistics 2018-09-10 Val Andrei Fajardo , David Findlay , Roshanak Houmanfar , Charu Jaiswal , Jiaxi Liang , Honglei Xie

Image Quality of MRI brain scans is strongly influenced by within scanner head movements and the resulting image artifacts alter derived measures like brain volume and cortical thickness. Automated image quality assessment is key to…

Image and Video Processing · Electrical Eng. & Systems 2024-11-05 Felix Hoffstaedter , Nicolás Nieto , Simon B. Eickhoff , Kaustubh R. Patil

This paper considers fixed effects (FE) estimation for linear panel data models under possible model misspecification when both the number of individuals, $n$, and the number of time periods, $T$, are large. We first clarify the probability…

Statistics Theory · Mathematics 2014-03-12 Antonio F. Galvao , Kengo Kato

In latent variable models the parameter estimation can be implemented by using the joint or the marginal likelihood, based on independence or conditional independence assumptions. The same dilemma occurs within the Bayesian framework with…

Computation · Statistics 2014-09-18 Silia Vitoratou , Ioannis Ntzoufras , Irini Moustaki

Epidemiologic studies often evaluate the association between an exposure and an event risk. When time-varying, exposure updates usually occur at discrete visits although changes are in continuous time and survival models require values to…

Data analysis based on information from several sources is common in economic and biomedical studies. This setting is often referred to as the data fusion problem, which differs from traditional missing data problems since no complete data…

Methodology · Statistics 2022-04-07 Wei Li , Shanshan Luo , Wangli Xu

Binwise Variance Scaling (BVS) has recently been proposed as a post hoc recalibration method for prediction uncertainties of machine learning regression problems that is able of more efficient corrections than uniform variance (or…

Machine Learning · Statistics 2023-10-25 Pascal Pernot