English
Related papers

Related papers: Agnostic Sample Compression Schemes for Regression

200 papers

Prediction, in regression and classification, is one of the main aims in modern data science. When the number of predictors is large, a common first step is to reduce the dimension of the data. Sufficient dimension reduction (SDR) is a well…

Methodology · Statistics 2023-06-21 Liliana Forzani , Daniela Rodriguez , Mariela Sued

In compressed sensing, in order to recover a sparse or nearly sparse vector from possibly noisy measurements, the most popular approach is $\ell_1$-norm minimization. Upper bounds for the $\ell_2$- norm of the error between the true and…

Machine Learning · Statistics 2015-12-31 M. Eren Ahsen , M. Vidyasagar

We develop a model-free theory of general types of parametric regression for iid observations. The theory replaces the parameters of parametric models with statistical functionals, to be called "regression functionals'', defined on large…

Statistics Theory · Mathematics 2019-07-09 Andreas Buja , Lawrence Brown , Arun Kumar Kuchibhotla , Richard Berk , Ed George , Linda Zhao

In the context of goal-oriented communications, this paper addresses the achievable rate versus generalization error region of a learning task applied on compressed data. The study focuses on the distributed setup where a source is…

Information Theory · Computer Science 2024-07-10 Jiahui Wei , Philippe Mary , Elsa Dupraz

We investigate Learning from Label Proportions (LLP), a partial information setting where examples in a training set are grouped into bags, and only aggregate label values in each bag are available. Despite the partial observability, the…

Machine Learning · Computer Science 2025-06-02 Robert Busa-Fekete , Travis Dick , Claudio Gentile , Haim Kaplan , Tomer Koren , Uri Stemmer

Hypothesis testing procedures are developed to assess linear operator constraints in function-on-scalar regression when incomplete functional responses are observed. The approach enables statistical inferences about the shape and other…

Methodology · Statistics 2022-12-06 Yeonjoo Park , Kyunghee Han , Douglas G. Simpson

A long-standing sample compression conjecture asks to linearly bound the size of the optimal sample compression schemes by the Vapnik-Chervonenkis (VC) dimension of an arbitrary class. In this paper, we explore the rich metric and…

Combinatorics · Mathematics 2024-03-08 Tilen Marc

We consider a family of infinite dimensional product measures with tails between Gaussian and exponential, which we call $p$-exponential measures. We study their measure-theoretic properties and in particular their concentration. Our…

Statistics Theory · Mathematics 2020-10-09 Sergios Agapiou , Masoumeh Dashti , Tapio Helin

Objective: Provide guidance on sample size considerations for developing predictive models by empirically establishing the adequate sample size, which balances the competing objectives of improving model performance and reducing model…

Applications · Statistics 2024-07-25 Luis H. John , Jan A. Kors , Jenna M. Reps , Patrick B. Ryan , Peter R. Rijnbeek

In this paper we develop a general theory of compressed sensing for analog signals, in close similarity to prior results for vectors in finite dimensional spaces that are sparse in a given orthonormal basis. The signals are modeled by…

Functional Analysis · Mathematics 2018-03-13 Bernard G. Bodmann , Axel Flinth , Gitta Kutyniok

We study the performance of empirical risk minimization on the $p$-norm linear regression problem for $p \in (1, \infty)$. We show that, in the realizable case, under no moment assumptions, and up to a distribution-dependent constant,…

Statistics Theory · Mathematics 2024-06-19 Ayoub El Hanchi , Murat A. Erdogdu

An important problem in space-time adaptive detection is the estimation of the large p-by-p interference covariance matrix from training signals. When the number of training signals n is greater than 2p, existing estimators are generally…

Signal Processing · Electrical Eng. & Systems 2021-07-26 Benjamin D. Robinson , Robert Malinas , Alfred O. Hero

Consider the communication-constrained estimation of discrete distributions under $\ell^p$ losses, where each distributed terminal holds multiple independent samples and uses limited number of bits to describe the samples. We obtain the…

Machine Learning · Computer Science 2024-11-11 Deheng Yuan , Tao Guo , Zhongyi Huang

We consider the problem of learning a target function corresponding to a deep, extensive-width, non-linear neural network with random Gaussian weights. We consider the asymptotic limit where the number of samples, the input dimension and…

Machine Learning · Statistics 2023-09-07 Hugo Cui , Florent Krzakala , Lenka Zdeborová

We prove a new generalization bound that shows for any class of linear predictors in Gaussian space, the Rademacher complexity of the class and the training error under any continuous loss $\ell$ can control the test error under all Moreau…

Machine Learning · Statistics 2022-10-24 Lijia Zhou , Frederic Koehler , Pragya Sur , Danica J. Sutherland , Nathan Srebro

We consider nonparametric estimation of a regression function for a situation where precisely measured predictors are used to estimate the regression curve for coarsened, that is, less precise or contaminated predictors. Specifically, while…

Statistics Theory · Mathematics 2008-12-18 Aurore Delaigle , Peter Hall , Hans-Georg Müller

We analyze the convergence of compressive sensing based sampling techniques for the efficient evaluation of functionals of solutions for a class of high-dimensional, affine-parametric, linear operator equations which depend on possibly…

Numerical Analysis · Mathematics 2015-09-22 Holger Rauhut , Christoph Schwab

We study the problem of compression for the purpose of similarity identification, where similarity is measured by the mean square Euclidean distance between vectors. While the asymptotical fundamental limits of the problem - the minimal…

Information Theory · Computer Science 2014-05-13 Fabian Steiner , Steffen Dempfle , Amir Ingber , Tsachy Weissman

This article studies the achievable guarantees on the error rates of certain learning algorithms, with particular focus on refining logarithmic factors. Many of the results are based on a general technique for obtaining bounds on the error…

Machine Learning · Computer Science 2016-09-13 Steve Hanneke

Let $\mathbb{T}^d$ denote the $d$-dimensional torus. We consider the problem of optimally recovering a target function $f^*:\mathbb{T}^d\rightarrow \mathbb{C}$ from samples of its Fourier coefficients. We make classical smoothness…

Functional Analysis · Mathematics 2025-09-01 Jonathan W. Siegel