中文
相关论文

相关论文: Agnostic Sample Compression Schemes for Regression

200 篇论文

We present a framework for the theoretical analysis of ensembles of low-complexity empirical risk minimisers trained on independent random compressions of high-dimensional data. First we introduce a general distribution-dependent…

机器学习 · 计算机科学 2021-06-03 Henry W. J. Reeve , Ata Kaban

We show that the topes of a complex of oriented matroids (abbreviated COM) of VC-dimension $d$ admit a proper labeled sample compression scheme of size $d$. This considerably extends results of Moran and Warmuth on ample classes, of…

组合数学 · 数学 2023-04-21 Victor Chepoi , Kolja Knauer , Manon Philibert

We introduce a new and improved characterization of the label complexity of disagreement-based active learning, in which the leading quantity is the version space compression set size. This quantity is defined as the size of the smallest…

机器学习 · 计算机科学 2014-04-08 Yair Wiener , Steve Hanneke , Ran El-Yaniv

We study the approximation of arbitrary distributions $P$ on $d$-dimensional space by distributions with log-concave density. Approximation means minimizing a Kullback--Leibler-type functional. We show that such an approximation exists if…

统计理论 · 数学 2011-10-17 Lutz Duembgen , Richard Samworth , Dominic Schuhmacher

We study a new class of codes for lossy compression with the squared-error distortion criterion, designed using the statistical framework of high-dimensional linear regression. Codewords are linear combinations of subsets of columns of a…

信息论 · 计算机科学 2015-12-21 Ramji Venkataramanan , Antony Joseph , Sekhar Tatikonda

A significant hurdle for analyzing large sample data is the lack of effective statistical computing and inference methods. An emerging powerful approach for analyzing large sample data is subsampling, by which one takes a random subsample…

统计方法学 · 统计学 2015-11-24 Rong Zhu , Ping Ma , Michael W. Mahoney , Bin Yu

Evaluating the statistical dimension is a common tool to determine the asymptotic phase transition in compressed sensing problems with Gaussian ensemble. Unfortunately, the exact evaluation of the statistical dimension is very difficult and…

信息论 · 计算机科学 2019-06-06 Sajad Daei , Farzan Haddadi , Arash Amini , Martin Lotz

It was proved in 1998 by Ben-David and Litman that a concept space has a sample compression scheme of size d if and only if every finite subspace has a sample compression scheme of size d. In the compactness theorem, measurability of the…

机器学习 · 统计学 2015-03-20 Damjan Kalajdzievski

We study the problem of selecting features associated with extreme values in high dimensional linear regression. Normally, in linear modeling problems, the presence of abnormal extreme values or outliers is considered an anomaly which…

统计方法学 · 统计学 2021-06-16 Andersen Chang , Minjie Wang , Genevera Allen

Accurate coresets are a weighted subset of the original dataset, ensuring a model trained on the accurate coreset maintains the same level of accuracy as a model trained on the full dataset. Primarily, these coresets have been studied for a…

机器学习 · 计算机科学 2024-12-31 Sanskar Ranjan , Supratim Shit

This work investigates the problem of signal recovery from undersampled noisy sub-Gaussian measurements under the assumption of a synthesis-based sparsity model. Solving the $\ell^1$-synthesis basis pursuit allows for a simultaneous…

信息论 · 计算机科学 2020-04-16 Maximilian März , Claire Boyer , Jonas Kahn , Pierre Weiss

We study the recovery results of $\ell_p$-constrained compressive sensing (CS) with $p\geq 1$ via robust width property and determine conditions on the number of measurements for standard Gaussian matrices under which the property holds…

信息论 · 计算机科学 2017-08-28 Zhiyong Zhou , Jun Yu

We analyze the properties of adversarial training for learning adversarially robust halfspaces in the presence of agnostic label noise. Denoting $\mathsf{OPT}_{p,r}$ as the best robust classification error achieved by a halfspace that is…

机器学习 · 计算机科学 2021-04-20 Difan Zou , Spencer Frei , Quanquan Gu

We address the problem of learning an unknown smooth function and its derivatives from noisy pointwise evaluations under the supremum norm. While classical nonparametric regression provides a strong theoretical foundation, traditional…

机器学习 · 计算机科学 2026-03-10 Davide Maran , Marcello Restelli

In large scale machine learning, random sampling is a popular way to approximate datasets by a small representative subset of examples. In particular, sensitivity sampling is an intensely studied technique which provides provable guarantees…

数据结构与算法 · 计算机科学 2024-01-04 David P. Woodruff , Taisuke Yasuda

We provide fast algorithms for overconstrained $\ell_p$ regression and related problems: for an $n\times d$ input matrix $A$ and vector $b\in\mathbb{R}^n$, in $O(nd\log n)$ time we reduce the problem $\min_{x\in\mathbb{R}^d} \|Ax-b\|_p$ to…

数据结构与算法 · 计算机科学 2014-04-08 Kenneth L. Clarkson , Petros Drineas , Malik Magdon-Ismail , Michael W. Mahoney , Xiangrui Meng , David P. Woodruff

This paper studies several aspects of signal reconstruction of sampled data in spaces of bandlimited functions. In the first part, signal spaces are characterized in which the classical sampling series uniformly converge, and we investigate…

信息论 · 计算机科学 2014-10-23 Holger Boche , Volker Pohl

We address the problem of nonparametric estimation of characteristics for stationary and ergodic time series. We consider finite-alphabet time series and real-valued ones and the following four problems: i) estimation of the (limiting)…

信息论 · 计算机科学 2007-11-01 Boris Ryabko

The problem of variable-rate lossless data compression is considered, for codes with and without prefix constraints. Sharp bounds are derived for the best achievable compression rate of memoryless sources, when the excess-rate probability…

信息论 · 计算机科学 2025-11-13 Andreas Theocharous , Lampros Gavalakis , Ioannis Kontoyiannis

We present a new general-purpose algorithm for learning classes of $[0,1]$-valued functions in a generalization of the prediction model, and prove a general upper bound on the expected absolute error of this algorithm in terms of a…

机器学习 · 计算机科学 2023-04-25 Peter L. Bartlett , Philip M. Long