English
Related papers

Related papers: Dimension-free uniform concentration bound for log…

200 papers

For the last two decades, high-dimensional data and methods have proliferated throughout the literature. Yet, the classical technique of linear regression has not lost its usefulness in applications. In fact, many high-dimensional…

Statistics Theory · Mathematics 2021-05-18 Arun Kumar Kuchibhotla , Lawrence D. Brown , Andreas Buja , Edward I. George , Linda Zhao

Order statistics theory is applied in this paper to probabilistic robust control theory to compute the minimum sample size needed to come up with a reliable estimate of an uncertain quantity under continuity assumption of the related…

Optimization and Control · Mathematics 2008-05-13 Xinjia Chen , Kemin Zhou

We consider high-dimensional binary classification by sparse logistic regression. We propose a model/feature selection procedure based on penalized maximum likelihood with a complexity penalty on the model size and derive the non-asymptotic…

Statistics Theory · Mathematics 2018-11-20 Felix Abramovich , Vadim Grinshtein

Under the assumption that the distribution of a nonnegative random variable $X$ admits a bounded coupling with its size biased version, we prove simple and strong concentration bounds. In particular the upper tail probability is shown to…

Probability · Mathematics 2014-07-15 Richard Arratia , Peter Baxendale

This short paper concerns a diffusive logistic equation with the heterogeneous environment and a free boundary, which is formulated to study the spread of an invasive species, where the free boundary represents the expanding front. A…

Analysis of PDEs · Mathematics 2014-06-25 Mingxin Wang

In this work, we introduce a modified (rescaled) likelihood for imbalanced logistic regression. This new approach makes easier the use of exponential priors and the computation of lasso regularization path. Precisely, we study a limiting…

Methodology · Statistics 2018-04-19 Vincent Runge

This note extends the results of classical parametric statistics like Fisher and Wilks theorem to modern setups with a high or infinite parameter dimension, limited sample size, and possible model misspecification. We consider a special…

Statistics Theory · Mathematics 2025-06-09 Vladimir Spokoiny

This paper formalizes a latent variable inference problem we call {\em supervised pattern discovery}, the goal of which is to find sets of observations that belong to a single ``pattern.'' We discuss two versions of the problem and prove…

Machine Learning · Statistics 2014-02-10 Jonathan H. Huggins , Cynthia Rudin

We propose an extensive analysis of the behavior of majority votes in binary classification. In particular, we introduce a risk bound for majority votes, called the C-bound, that takes into account the average quality of the voters and…

Machine Learning · Statistics 2015-07-30 Pascal Germain , Alexandre Lacasse , François Laviolette , Mario Marchand , Jean-Francis Roy

While standard statistical inference techniques and machine learning generalization bounds assume that tests are run on data selected independently of the hypotheses, practical data analysis and machine learning are usually iterative and…

Signal Processing · Electrical Eng. & Systems 2019-10-09 Lorenzo De Stefani , Eli Upfal

Constrained diffusions in convex polyhedral domains with a general oblique reflection field, and with a diffusion coefficient scaled by a small parameter, are considered. Using an interior Dirichlet heat kernel lower bound estimate for…

Probability · Mathematics 2013-08-19 Amarjit Budhiraja , Zhen-Qing Chen

Standard Bayesian learning is known to have suboptimal generalization capabilities under misspecification and in the presence of outliers. PAC-Bayes theory demonstrates that the free energy criterion minimized by Bayesian learning is a…

Machine Learning · Computer Science 2023-04-25 Matteo Zecchin , Sangwoo Park , Osvaldo Simeone , Marios Kountouris , David Gesbert

In many contemporary statistical and machine learning methods, one needs to optimize an objective function that depends on the discrepancy between two probability distributions. The discrepancy can be referred to as a metric for…

Machine Learning · Computer Science 2025-02-11 Yijin Ni , Xiaoming Huo

In cluster-specific studies, ordinary logistic regression and conditional logistic regression for binary outcomes provide maximum likelihood estimator (MLE) and conditional maximum likelihood estimator (CMLE), respectively. In this paper,…

Statistics Theory · Mathematics 2020-05-14 Zhulin He , Yuyuan Ouyang

We study the behavior of linear discriminant functions for binary classification in the infinite-imbalance limit, where the sample size of one class grows without bound while the sample size of the other remains fixed. The coefficients of…

Machine Learning · Statistics 2023-05-15 Paul Glasserman , Mike Li

Offset Rademacher complexities have been shown to provide tight upper bounds for the square loss in a broad class of problems including improper statistical learning and online learning. We show that the offset complexity can be generalized…

Machine Learning · Statistics 2021-10-27 Suhas Vijaykumar

The fat-shattering dimension characterizes the uniform convergence property of real-valued functions. The state-of-the-art upper bounds feature a multiplicative squared logarithmic factor on the sample complexity, leaving an open gap with…

Machine Learning · Computer Science 2023-07-14 Roberto Colomboni , Emmanuel Esposito , Andrea Paudice

We obtain error approximation bounds between expected suprema of canonical processes that are generated by random vectors with independent coordinates and expected suprema of Gaussian processes. In particular, we obtain a sharper proximity…

Probability · Mathematics 2024-11-06 Shivam Sharma

We present a set of high-probability inequalities that control the concentration of weighted averages of multiple (possibly uncountably many) simultaneously evolving and interdependent martingales. Our results extend the PAC-Bayesian…

Machine Learning · Computer Science 2012-07-31 Yevgeny Seldin , François Laviolette , Nicolò Cesa-Bianchi , John Shawe-Taylor , Peter Auer

This article deals with the generalization performance of margin multi-category classifiers, when minimal learnability hypotheses are made. In that context, the derivation of a guaranteed risk is based on the handling of capacity measures…

Machine Learning · Computer Science 2020-09-17 Yann Guermeur