中文
相关论文

相关论文: Model Selection for High-Dimensional Regression un…

200 篇论文

The central problem we address in this work is estimation of the parameter support set S, the set of indices corresponding to nonzero parameters, in the context of a sparse parametric likelihood model for discrete multivariate time series.…

统计方法学 · 统计学 2025-03-07 Trevor D. Ruiz , Sharmodeep Bhattacharyya , Sarah C. Emerson

Graphical LASSO (GLASSO) is a widely used method for estimating sparse precision matrices and learning undirected graphical models in high-dimensional settings. Because GLASSO penalizes entries of the precision matrix directly, however, it…

Sparse regularization such as $\ell_1$ regularization is a quite powerful and widely used strategy for high dimensional learning problems. The effectiveness of sparse regularization has been supported practically and theoretically by…

机器学习 · 统计学 2018-02-23 Masaaki Takada , Taiji Suzuki , Hironori Fujisawa

We study the theoretical properties of the fused lasso procedure originally proposed by \cite{tibshirani2005sparsity} in the context of a linear regression model in which the regression coefficient are totally ordered and assumed to be…

统计理论 · 数学 2023-06-28 Fan Wang , Oscar Hernan Madrid Padilla , Yi Yu , Alessandro Rinaldo

We present a new and general method of weighted least square univariate regression where the dependent variable is expanded as a series of suitably chosen functions of the independent variables. Each term of the series is obtained by an…

数值分析 · 数学 2021-03-26 Nilotpal Kanti Sinha

Variable selection for high-dimensional, highly correlated data has long been a challenging problem, often yielding unstable and unreliable models. We propose a resample-aggregate framework that exploits diffusion models' ability to…

统计方法学 · 统计学 2025-08-20 Minjie Wang , Xiaotong Shen , Wei Pan

In this paper, we investigate seemingly unrelated regression (SUR) models that allow the number of equations (N) to be large, and to be comparable to the number of the observations in each equation (T). It is well known in the literature…

计量经济学 · 经济学 2018-11-15 Lidan Tan , Khai X. Chiong , Hyungsik Roger Moon

Large-scale empirical data, the sample size and the dimension are high, often exhibit various characteristics. For example, the noise term follows unknown distributions or the model is very sparse that the number of critical variables is…

统计理论 · 数学 2018-06-18 Yuehan Yang , Hu Yang

Gaussian Graphical Models (GGMs) are popular tools for studying network structures. However, many modern applications such as gene network discovery and social interactions analysis often involve high-dimensional noisy data with outliers or…

机器学习 · 统计学 2015-10-30 Eunho Yang , Aurélie C. Lozano

This paper is concerned with inference about low-dimensional components of a high-dimensional parameter vector $\beta^0$ which is identified through instrumental variables. We allow for eigenvalues of the expected outer product of included…

计量经济学 · 经济学 2020-08-05 Christoph Breunig , Enno Mammen , Anna Simoni

Despite impressive performance as evaluated on i.i.d. holdout data, deep neural networks depend heavily on superficial statistics of the training data and are liable to break under distribution shift. For example, subtle changes to the…

计算机视觉与模式识别 · 计算机科学 2019-03-18 Haohan Wang , Zexue He , Zachary C. Lipton , Eric P. Xing

We theoretically analyze the model selection consistency of least absolute shrinkage and selection operator (Lasso), both with and without post-thresholding, for high-dimensional Ising models. For random regular (RR) graphs of size $p$ with…

机器学习 · 统计学 2023-02-20 Xiangming Meng , Tomoyuki Obuchi , Yoshiyuki Kabashima

We introduce a computationally effective algorithm for a linear model selection consisting of three steps: screening--ordering--selection (SOS). Screening of predictors is based on the thresholded Lasso that is l_1 penalized least squares.…

机器学习 · 统计学 2013-10-24 Piotr Pokarowski , Jan Mielniczuk

We propose a nonparametric procedure to achieve fast inference in generative graphical models when the number of latent states is very large. The approach is based on iterative latent variable preselection, where we alternate between…

机器学习 · 统计学 2021-10-06 Jacquelyn A. Shelton , Jan Gasthaus , Zhenwen Dai , Joerg Luecke , Arthur Gretton

Consider a regression model with infinitely many parameters and time series errors. We are interested in choosing weights for averaging across generalized least squares (GLS) estimators obtained from a set of approximating models. However,…

统计理论 · 数学 2016-10-05 Tzu-Chang F. Cheng , Ching-Kang Ing , Shu-Hui Yu

Consider the problem of estimating the mean of a Gaussian random vector when the mean vector is assumed to be in a given convex set. The most natural solution is to take the Euclidean projection of the data vector on to this convex set; in…

统计理论 · 数学 2014-11-21 Sourav Chatterjee

The lasso has been studied extensively as a tool for estimating the coefficient vector in the high-dimensional linear model; however, considerably less is known about estimating the error variance in this context. In this paper, we propose…

统计方法学 · 统计学 2019-07-22 Guo Yu , Jacob Bien

In many applications, particularly in the natural sciences, the available high-dimensional set of features may contain variables that are not correlated with the response under consideration. Such irrelevant features can, in certain cases,…

统计理论 · 数学 2025-07-28 Gianluca Finocchio , Tatyana Krivobokova

Nonparametric regression for massive numbers of samples (n) and features (p) is an increasingly important problem. In big n settings, a common strategy is to partition the feature space, and then separately apply simple models to each…

机器学习 · 统计学 2014-06-10 Rajarshi Guhaniyogi , David B. Dunson

The IBOSS approach proposed by Wang et al. (2019) selects the most informative subset of n points. It assumes that the ordinary least squares method is used and requires that the number of variables, p, is not large. However, in many…

统计方法学 · 统计学 2024-01-23 Xin Wang , Min Yang , William Li