中文
相关论文

相关论文: Covariate Selection Based on a Assumpton-free Appr…

200 篇论文

We consider the problem of estimating the conditional mean of a real Gaussian variable $\nolinebreak Y=\sum_{i=1}^p\nolinebreak\theta_iX_i+\nolinebreak \epsilon$ where the vector of the covariates $(X_i)_{1\leq i\leq p}$ follows a joint…

统计理论 · 数学 2009-04-28 Nicolas Verzelen

We study the task of noiseless linear regression under Gaussian covariates in the presence of additive oblivious contamination. Specifically, we are given i.i.d.\ samples from a distribution $(x, y)$ on $\mathbb{R}^d \times \mathbb{R}$ with…

数据结构与算法 · 计算机科学 2025-10-14 Ilias Diakonikolas , Chao Gao , Daniel M. Kane , John Lafferty , Ankit Pensia

Sparse linear regression is a central problem in high-dimensional statistics. We study the correlated random design setting, where the covariates are drawn from a multivariate Gaussian $N(0,\Sigma)$, and we seek an estimator with small…

数据结构与算法 · 计算机科学 2023-05-29 Jonathan Kelner , Frederic Koehler , Raghu Meka , Dhruv Rohatgi

We introduce a novel Bayesian approach for variable selection using Gaussian process regression, which is crucial for enhancing interpretability and model regularization. Our method employs nearest neighbor Gaussian processes, serving as…

In this paper we study the asymptotics of linear regression in settings with non-Gaussian covariates where the covariates exhibit a linear dependency structure, departing from the standard assumption of independence. We model the covariates…

机器学习 · 统计学 2024-12-10 Behrad Moniri , Hamed Hassani

In the high-dimensional regression model a response variable is linearly related to $p$ covariates, but the sample size $n$ is smaller than $p$. We assume that only a small subset of covariates is `active' (i.e., the corresponding…

统计理论 · 数学 2013-05-03 Adel Javanmard , Andrea Montanari

Recent work has focused on the problem of conducting linear regression when the number of covariates is very large, potentially greater than the sample size. To facilitate this, one useful tool is to assume that the model can be well…

统计方法学 · 统计学 2011-11-21 Zhou Fang

It is well-known that the statistical performance of Lasso can suffer significantly when the covariates of interest have strong correlations. In particular, the prediction error of Lasso becomes much worse than computationally inefficient…

机器学习 · 统计学 2024-02-26 Jonathan Kelner , Frederic Koehler , Raghu Meka , Dhruv Rohatgi

Variable selection comprises an important step in many modern statistical inference procedures. In the regression setting, when estimators cannot shrink irrelevant signals to zero, covariates without relationships to the response often…

统计理论 · 数学 2025-03-28 Ka Long Keith Ho , Hien Duy Nguyen

Covariate balancing is a popular technique for controlling confounding in observational studies. It finds weights for the treatment group which are close to uniform, but make the group's covariate means (approximately) equal to those of the…

统计方法学 · 统计学 2025-03-07 Shiva Kaul , Min-Gyu Kim

Complex-valued signals are used in the modeling of many systems in engineering and science, hence being of fundamental interest. Often, random complex-valued signals are considered to be proper. A proper complex random variable or process…

Subsampling is an efficient method to deal with massive data. In this paper, we investigate the optimal subsampling for linear quantile regression when the covariates are functions. The asymptotic distribution of the subsampling estimator…

数值分析 · 数学 2022-05-06 Qian Yan , Hanyu Li , Chengmei Niu

Given a Gaussian Markov random field, we consider the problem of selecting a subset of variables to observe which minimizes the total expected squared prediction error of the unobserved variables. We first show that finding an exact…

机器学习 · 计算机科学 2012-09-27 Satyaki Mahalanabis , Daniel Stefankovic

Gaussian process regression is a powerful Bayesian nonlinear regression method. Recent research has enabled the capture of many types of observations using non-Gaussian likelihoods. To deal with various tasks in spatial modeling, we benefit…

机器学习 · 统计学 2025-08-26 Yuta Shikuri

Adjusting for (baseline) covariates with working regression models becomes standard practice in the analysis of randomized clinical trials (RCT). When the dimension $p$ of the covariates is large relative to the sample size $n$,…

统计方法学 · 统计学 2025-12-24 Yujia Gu , Lin Liu , Wei Ma

We propose a computationally intensive method, the random lasso method, for variable selection in linear models. The method consists of two major steps. In step 1, the lasso method is applied to many bootstrap samples, each using a set of…

应用统计 · 统计学 2011-04-19 Sijian Wang , Bin Nan , Saharon Rosset , Ji Zhu

We consider the estimation of regression models on strata defined using a categorical covariate, in order to identify interactions between this categorical covariate and the other predictors. A basic approach requires the choice of a…

统计理论 · 数学 2016-11-09 Edouard Ollier , Vivian Viallon

Suppose that a data analyst wishes to report the results of a least squares linear regression only if the overall null hypothesis, $H_0^{1:p}: \beta_1= \beta_2 = \ldots = \beta_p=0$, is rejected. This practice, which we refer to as…

统计方法学 · 统计学 2026-05-12 Olivia McGough , Daniela Witten , Daniel Kessler

Standard penalized methods of variable selection and parameter estimation rely on the magnitude of coefficient estimates to decide which variables to include in the final model. However, coefficient estimates are unreliable when the design…

统计方法学 · 统计学 2018-02-13 Jonathan P Williams , Jan Hannig

We consider a Bayesian approach to model selection in Gaussian linear regression, where the number of predictors might be much larger than the number of observations. From a frequentist view, the proposed procedure results in the penalized…

统计理论 · 数学 2010-09-14 Felix Abramovich , Vadim Grinshtein