English
Related papers

Related papers: Dimension-free deterministic equivalents and scali…

200 papers

Kernel methods form a powerful, versatile, and theoretically-grounded unifying framework to solve nonlinear problems in signal processing and machine learning. The standard approach relies on the kernel trick to perform pairwise evaluations…

Machine Learning · Computer Science 2019-12-11 Kan Li , Jose C. Principe

Feature selection has evolved to be an important step in several machine learning paradigms. In domains like bio-informatics and text classification which involve data of high dimensions, feature selection can help in drastically reducing…

Machine Learning · Computer Science 2019-04-23 Nand Sharma , Prathamesh Verlekar , Rehab Ashary , Sui Zhiquan

Our main focus is on the generalization bound, which serves as an upper limit for the generalization error. Our analysis delves into regression and classification tasks separately to ensure a thorough examination. We assume the target…

Machine Learning · Statistics 2024-07-30 Wen-Liang Hwang

We consider the approximation in the reaction-diffusion norm with continuous finite elements and prove that the best error is equivalent to a sum of the local best errors on pairs of elements. The equivalence constants do not depend on the…

Numerical Analysis · Mathematics 2018-03-07 Francesca Tantardini , Andreas Veeser , R"udiger Verf"urth

We study the asymmetric matrix factorization problem under a natural nonconvex formulation with arbitrary overparametrization. The model-free setting is considered, with minimal assumption on the rank or singular values of the observed…

Machine Learning · Computer Science 2023-08-22 Liwei Jiang , Yudong Chen , Lijun Ding

Popularly used eigendecomposition-based criteria such as BIC type, ratio estimation and principal component-based criterion often underdetermine model dimensionality for regressions or the number of factors for factor models. This…

Statistics Theory · Mathematics 2016-08-17 Xuehu Zhu , Tao Wang , Lixing Zhu

The goal of supervised representation learning is to construct effective data representations for prediction. Among all the characteristics of an ideal nonparametric representation of high-dimensional complex data, sufficiency, low…

Machine Learning · Computer Science 2022-09-02 Jian Huang , Yuling Jiao , Xu Liao , Jin Liu , Zhou Yu

We study the relationship between gradient-based optimization of parametric models (e.g., neural networks) and optimization of linear combinations of random features. Our main result shows that if a parametric model can be learned using…

Machine Learning · Computer Science 2025-05-16 Ari Karchmer , Eran Malach

It has been observed that the performances of many high-dimensional estimation problems are universal with respect to underlying sensing (or design) matrices. Specifically, matrices with markedly different constructions seem to achieve…

Information Theory · Computer Science 2023-07-24 Rishabh Dudeja , Subhabrata Sen , Yue M. Lu

We develop some graph-based tests for spherical symmetry of a multivariate distribution using a method based on data augmentation. These tests are constructed using a new notion of signs and ranks that are computed along a path obtained by…

Statistics Theory · Mathematics 2024-12-10 Bilol Banerjee , Anil K. Ghosh

A fundamental problem in multivariate analysis is testing general linear hypotheses for regression coefficients in a multivariate linear model. This framework encompasses a wide range of well-studied tasks, including MANOVA, joint…

Methodology · Statistics 2025-07-09 Haoran Li

This paper studies the validity of nonparametric tests used in the regression discontinuity design. The null hypothesis of interest is that the average treatment effect at the threshold in the so-called sharp design equals a pre-specified…

Methodology · Statistics 2016-11-16 Vishal Kamat

Recent random-forest (RF)-based image super-resolution approaches inherit some properties from dictionary-learning-based algorithms, but the effectiveness of the properties in RF is overlooked in the literature. In this paper, we present a…

Computer Vision and Pattern Recognition · Computer Science 2017-12-15 Hailiang Li , Kin-Man Lam , Miaohui Wang

High-dimensional data is commonly encountered in numerous data analysis tasks. Feature selection techniques aim to identify the most representative features from the original high-dimensional data. Due to the absence of class label…

Machine Learning · Computer Science 2024-10-29 Yunhui Liang , Jianwen Gan , Yan Chen , Peng Zhou , Liang Du

This manuscript studies statistical properties of linear classifiers obtained through minimization of an unregularized convex risk over a finite sample. Although the results are explicitly finite-dimensional, inputs may be passed through…

Machine Learning · Computer Science 2012-06-15 Matus Telgarsky

We study high-dimensional, ridge-regularized logistic regression in a setting in which the covariates may be missing or corrupted by additive noise. When both the covariates and the additive corruptions are independent and normally…

Statistics Theory · Mathematics 2024-10-03 Kabir Aladin Verchand , Andrea Montanari

Prediction, in regression and classification, is one of the main aims in modern data science. When the number of predictors is large, a common first step is to reduce the dimension of the data. Sufficient dimension reduction (SDR) is a well…

Methodology · Statistics 2023-06-21 Liliana Forzani , Daniela Rodriguez , Mariela Sued

We consider the overfitting behavior of minimum norm interpolating solutions of Gaussian kernel ridge regression (i.e. kernel ridgeless regression), when the bandwidth or input dimension varies with the sample size. For fixed dimensions, we…

Machine Learning · Computer Science 2024-09-09 Marko Medvedev , Gal Vardi , Nathan Srebro

Recently, there has been remarkable progress in reinforcement learning (RL) with general function approximation. However, all these works only provide regret or sample complexity guarantees. It is still an open question if one can achieve…

Machine Learning · Computer Science 2023-05-16 Yue Wu , Jiafan He , Quanquan Gu

Although a majority of the theoretical literature in high-dimensional statistics has focused on settings which involve fully-observed data, settings with missing values and corruptions are common in practice. We consider the problems of…

Machine Learning · Statistics 2017-11-06 Yining Wang , Jialei Wang , Sivaraman Balakrishnan , Aarti Singh