English
Related papers

Related papers: Distribution-free tests for lossless feature selec…

200 papers

Variable selection comprises an important step in many modern statistical inference procedures. In the regression setting, when estimators cannot shrink irrelevant signals to zero, covariates without relationships to the response often…

Statistics Theory · Mathematics 2025-03-28 Ka Long Keith Ho , Hien Duy Nguyen

Feature screening approaches are effective in selecting active features from data with ultrahigh dimensionality and increasing complexity; however, the majority of existing feature screening approaches are either restricted to a univariate…

Methodology · Statistics 2023-05-09 Shaofei Zhao , Guifang Fu

We study the problem of out-of-sample risk estimation in the high dimensional regime where both the sample size $n$ and number of features $p$ are large, and $n/p$ can be less than one. Extensive empirical evidence confirms the accuracy of…

Machine Learning · Statistics 2020-03-05 Kamiar Rahnama Rad , Wenda Zhou , Arian Maleki

We investigate fast methods that allow to quickly eliminate variables (features) in supervised learning problems involving a convex loss function and a $l_1$-norm penalty, leading to a potentially substantial reduction in the number of…

Machine Learning · Computer Science 2010-10-28 Laurent El Ghaoui , Vivian Viallon , Tarek Rabbani

Multivariate extreme value statistical analysis is concerned with observations on several variables which are thought to possess some degree of tail-dependence. In areas such as the modeling of financial and insurance risks, or as the…

Applications · Statistics 2014-12-31 Alexis Bienvenüe , Christian Y. Robert

We study the low rank regression problem $\my = M\mx + \epsilon$, where $\mx$ and $\my$ are $d_1$ and $d_2$ dimensional vectors respectively. We consider the extreme high-dimensional setting where the number of observations $n$ is less than…

Data Structures and Algorithms · Computer Science 2020-10-27 Qiong Wu , Felix Ming Fai Wong , Zhenming Liu , Yanhua Li , Varun Kanade

We study density estimation for classes of shift-invariant distributions over $\mathbb{R}^d$. A multidimensional distribution is "shift-invariant" if, roughly speaking, it is close in total variation distance to a small shift of it in any…

Machine Learning · Computer Science 2018-11-12 Anindya De , Philip M. Long , Rocco A. Servedio

We propose a new sufficient dimension reduction approach designed deliberately for high-dimensional classification. This novel method is named maximal mean variance (MMV), inspired by the mean variance index first proposed by Cui, Li and…

Methodology · Statistics 2018-12-11 Xin Chen , Jingjing Wu , Zhigang Yao , Jia Zhang

Change point testing for high-dimensional data has attracted a lot of attention in statistics and machine learning owing to the emergence of high-dimensional data with structural breaks from many fields. In practice, when the dimension is…

Methodology · Statistics 2023-12-05 Hanjia Gao , Runmin Wang , Xiaofeng Shao

We conduct a non asymptotic study of the Cross Validation (CV) estimate of the generalization risk for learning algorithms dedicated to extreme regions of the covariates space. In this Extreme Value Analysis context, the risk function…

Statistics Theory · Mathematics 2024-09-12 Anass Aghbalou , Patrice Bertail , François Portier , Anne Sabourin

We consider the problem of predicting as well as the best linear combination of d given functions in least squares regression, and variants of this problem including constraints on the parameters of the linear combination. When the input…

Machine Learning · Statistics 2010-07-06 Jean-Yves Audibert , Olivier Catoni

This article provides, through theoretical analysis, an in-depth understanding of the classification performance of the empirical risk minimization framework, in both ridge-regularized and unregularized cases, when high dimensional data are…

Machine Learning · Statistics 2020-11-26 Xiaoyi Mai , Zhenyu Liao

We consider the following basic, and very broad, statistical problem: Given a known high-dimensional distribution ${\cal D}$ over $\mathbb{R}^n$ and a collection of data points in $\mathbb{R}^n$, distinguish between the two possibilities…

Computational Complexity · Computer Science 2024-11-25 Anindya De , Huan Li , Shivam Nadimpalli , Rocco A. Servedio

Sparse linear regression is a central problem in high-dimensional statistics. We study the correlated random design setting, where the covariates are drawn from a multivariate Gaussian $N(0,\Sigma)$, and we seek an estimator with small…

Data Structures and Algorithms · Computer Science 2023-05-29 Jonathan Kelner , Frederic Koehler , Raghu Meka , Dhruv Rohatgi

Suppose that the only available information in a multi-class problem are expert estimates of the conditional probabilities of occurrence for a set of binary features. The aim is to select a subset of features to be measured in subsequent…

Artificial Intelligence · Computer Science 2012-07-19 Ludmila Kuncheva , C. Whitaker , P. Cockcroft , Z. S. Hoare

In this paper we estimate the mean-variance portfolio in the high-dimensional case using the recent results from the theory of random matrices. We construct a linear shrinkage estimator which is distribution-free and is optimal in the sense…

Statistical Finance · Quantitative Finance 2023-04-19 Taras Bodnar , Yarema Okhrin , Nestor Parolya

Bayesian variable selection has gained much empirical success recently in a variety of applications when the number $K$ of explanatory variables $(x_1,...,x_K)$ is possibly much larger than the sample size $n$. For generalized linear…

Statistics Theory · Mathematics 2009-09-29 Wenxin Jiang

The goal of feature selection is to choose the optimal subset of features for a recognition task by evaluating the importance of each feature, thereby achieving effective dimensionality reduction. Currently, proposed feature selection…

Machine Learning · Computer Science 2024-02-27 Zhenxing Zhang , Jun Ge , Zheng Wei , Chunjie Zhou , Yilei Wang

Semi-supervised classification, where unlabeled data are massive but labeled data are limited, often arises in machine learning applications. We address this challenge under high-dimensional data by leveraging the manifold and cluster…

Machine Learning · Statistics 2026-04-28 Ruoxu Tan , Yiming Zang

Many high-dimensional hypothesis tests aim to globally examine marginal or low-dimensional features of a high-dimensional joint distribution, such as testing of mean vectors, covariance matrices and regression coefficients. This paper…

Statistics Theory · Mathematics 2020-02-04 Yinqiu He , Gongjun Xu , Chong Wu , Wei Pan
‹ Prev 1 8 9 10 Next ›