English
Related papers

Related papers: The effect of collinearity and sample size on line…

200 papers

Multicollinearity is relevant to many different fields where linear regression models are applied, and its existence may affect the analysis of ordinary least squares (OLS) estimators from both the numerical and statistical points of views.…

Methodology · Statistics 2020-05-06 Román Salmerón , Catalina García , José García

This paper shows that the degree of approximate multicollinearity in a linear regression model increases simply by including independent variables, even if these are not highly linearly related. In the current situation where it is…

Methodology · Statistics 2025-03-07 Román Salmerón Gómez , Catalina García García

The sophisticated and automated means of data collection used by an increasing number of institutions and companies leads to extremely large data sets. Subset selection in regression is essential when a huge number of covariates can…

Applications · Statistics 2013-04-22 Debbie J. Dupuis , Maria-Pia Victoria-Feser

A result from a standard linear model course is that the variance of the ordinary least squares (OLS) coefficient of a variable will never decrease when including additional covariates. The variance inflation factor (VIF) measures the…

Methodology · Statistics 2020-06-08 Peng Ding

Multicollinearity produces an inflation in the variance of the Ordinary Least Squares estimators due to the correlation between two or more independent variables (including the constant term). A widely applied solution is to estimate with…

Methodology · Statistics 2021-04-30 Román Salmerón Gómez , Catalina García García , José García Pérez

We consider the problem of variable selection in Bayesian multivariate linear regression models, involving multiple response and predictor variables, under multivariate normal errors. In the absence of a known covariance structure,…

Methodology · Statistics 2025-07-25 Joyee Ghosh , Xun Li

Building on key papers that were published in special issues of Biometrics in 1957 and 1982 we propose and develop a three-aspect system for evaluating the effect of fitting covariates in the analysis of designed experiments, in particular…

Methodology · Statistics 2025-08-08 Stephen Senn , Franz König , Martin Posch

Clinical prediction models are developed widely across medical disciplines. When predictors in such models are highly collinear, unexpected or spurious predictor-outcome associations may occur, thereby potentially reducing face-validity and…

In this paper, we address the problem of convergence of sequential variational inference filter (VIF) through the application of a robust variational objective and Hinf-norm based correction for a linear Gaussian system. As the dimension of…

Machine Learning · Statistics 2022-04-28 Niladri Das , Jed A. Duersch , Thomas A. Catanach

Regression calibration as developed by Rosner, Spiegelman and Willet is used to correct the bias in effect estimates due to measurement error in continuous exposures. The method involves two models: a measurement error model (MEM) relating…

Methodology · Statistics 2026-02-24 Wenze Tang , Donna Spiegelman , Xiaomei Liao , Molin Wang

The cos-max method is a little-known method of identifying collinearities. It is based on the cos-max transformation [4], which makes minimal adjustment to a set of vectors to create orthogonal components with a one-to-one correspondence…

Methodology · Statistics 2023-06-13 Zillur R. Shabuz , Paul H. Garthwaite

We consider the problem of learning linear prediction models with model misspecification bias. In such case, the collinearity among input variables may inflate the error of parameter estimation, resulting in instability of prediction…

Machine Learning · Computer Science 2019-12-02 Zheyan Shen , Peng Cui , Tong Zhang , Kun Kuang

Both classification and regression tasks are susceptible to the biased distribution of training data. However, existing approaches are focused on the class-imbalanced learning and cannot be applied to the problems of numerical regression…

Machine Learning · Computer Science 2021-09-15 Wentai Wu , Ligang He , Weiwei Lin

Regression analysis based on many covariates is becoming increasingly common. However, when the number of covariates $p$ is of the same order as the number of observations $n$, maximum likelihood regression becomes unreliable due to…

Methodology · Statistics 2023-09-06 Emanuele Massa , Marianne Jonker , Kit Roes , Anthony Coolen

To derive recommendations on how to analyze longitudinal data, we examined Type I error rates of Multilevel Linear Models (MLM) and repeated measures Analysis of Variance (rANOVA) using SAS and SPSS.We performed a simulation with the…

Applications · Statistics 2018-11-06 Nicolas Haverkamp , André Beauducel

In many machine learning problems, understanding variable importance is a central concern. Two common approaches are Permute-and-Predict (PaP), which randomly permutes a feature in a validation set, and Leave-One-Covariate-Out (LOCO), which…

Statistics Theory · Mathematics 2025-10-02 Kelvyn K. Bladen , D. Richard Cutler , Alan Wisler

Calibration is a vital aspect of the performance of risk prediction models, but research in the context of ordinal outcomes is scarce. This study compared calibration measures for risk models predicting a discrete ordinal outcome, and…

This study addresses a fundamental, yet overlooked, gap between standard theory and empirical modelling practices in the OLS regression model $\boldsymbol{y}=\boldsymbol{X\beta}+\boldsymbol{u}$ with collinearity. In fact, while an estimated…

Methodology · Statistics 2023-06-27 Takeaki Kariya , Hiroshi Kurata , Takaki Hayashi

Linear regression is a frequently used tool in statistics, however, its validity and interpretability relies on strong model assumptions. While robust estimates of the coefficients' covariance extend the validity of hypothesis tests and…

Methodology · Statistics 2015-04-23 Werner Brannath , Martin Scharpenberg

We develop a Coordinate Ascent Variational Inference (CAVI) algorithm for Bayesian Mixed Data Sampling (MIDAS) regression with linear weight parameterizations. The model separates impact coeffcients from weighting function parameters…

Machine Learning · Computer Science 2026-02-24 Luigi Simeone
‹ Prev 1 2 3 10 Next ›