English
Related papers

Related papers: Quantizing Heavy-tailed Data in Statistical Estima…

200 papers

In this paper, we provide novel optimal (or near optimal) convergence rates for a clipped version of the stochastic subgradient method. We consider nonsmooth convex problems over possibly unbounded domains, under heavy-tailed noise that…

Optimization and Control · Mathematics 2025-04-21 Daniela Angela Parletta , Andrea Paudice , Saverio Salzo

Dataset distillation creates a small distilled set that enables efficient training by capturing key information from the full dataset. While existing dataset distillation methods perform well on balanced datasets, they struggle under…

Computer Vision and Pattern Recognition · Computer Science 2025-12-16 Xiao Cui , Yulei Qin , Xinyue Li , Wengang Zhou , Hongsheng Li , Houqiang Li

In optimal covariance cleaning theory, minimizing the Frobenius norm between the true population covariance matrix and a rotational invariant estimator is a key step. This estimator can be obtained asymptotically for large covariance…

Information Theory · Computer Science 2023-05-01 Christian Bongiorno , Marco Berritta

Model compression has gained a lot of attention due to its ability to reduce hardware resource requirements significantly while maintaining accuracy of DNNs. Model compression is especially useful for memory-intensive recurrent neural…

Machine Learning · Computer Science 2018-05-30 Dongsoo Lee , Byeongwook Kim

Modern scientific instruments produce vast amounts of data, which can overwhelm the processing ability of computer systems. Lossy compression of data is an intriguing solution, but comes with its own drawbacks, such as potential signal…

We consider the problem of approximating the set of eigenvalues of the covariance matrix of a multivariate distribution (equivalently, the problem of approximating the "population spectrum"), given access to samples drawn from the…

Machine Learning · Computer Science 2017-07-18 Weihao Kong , Gregory Valiant

Compressed Sensing suggests that the required number of samples for reconstructing a signal can be greatly reduced if it is sparse in a known discrete basis, yet many real-world signals are sparse in a continuous dictionary. One example is…

Information Theory · Computer Science 2015-07-24 Yuanxin Li , Yuejie Chi

We present new estimators of the mean of a real valued random variable, based on PAC-Bayesian iterative truncation. We analyze the non-asymptotic minimax properties of the deviations of estimators for distributions having either a bounded…

Statistics Theory · Mathematics 2009-09-30 Olivier Catoni

We study tail estimation in Pareto-like settings for datasets with a high percentage of randomly right-censored data, and where some expert information on the tail index is available for the censored observations. This setting arises for…

Applications · Statistics 2019-11-13 Martin Bladt , Hansjoerg Albrecher , Jan Beirlant

We present optimal sample complexity estimates for one-bit compressed sensing problems in a realistic scenario: the procedure uses a structured matrix (a randomly sub-sampled circulant matrix) and is robust to analog pre-quantization noise…

Information Theory · Computer Science 2018-12-18 Sjoerd Dirksen , Shahar Mendelson

This paper investigates the asymptotic properties of quantile regression estimators in linear models, with a particular focus on polynomial regressors and robustness to heavy-tailed noise. Under independent and identically distributed…

Statistics Theory · Mathematics 2025-06-09 Saïd Maanan , Azzouz Dermoune , Ahmed El Ghini

This paper studies the computational and statistical aspects of quantile and pseudo-Huber tensor decomposition. The integrated investigation of computational and statistical issues of robust tensor decomposition poses challenges due to the…

Statistics Theory · Mathematics 2023-09-07 Yinan Shen , Dong Xia

As data volume grows extensively, data profiling helps to extract metadata of large-scale data. However, one kind of metadata, order statistics, is difficult to be computed because they are not mergeable or incremental. Thus, the limitation…

Data Structures and Algorithms · Computer Science 2020-06-29 Zhiwei Chen , Aoqian Zhang

In this paper we study recovery conditions of weighted $\ell_1$ minimization for signal reconstruction from compressed sensing measurements when partial support information is available. We show that if at least 50% of the (partial) support…

Information Theory · Computer Science 2011-07-26 Michael P. Friedlander , Hassan Mansour , Rayan Saab , Ozgur Yilmaz

We investigate the computational issues related to the memory size in the estimation of quadratic covariation, taking into account the specifics of financial ultra-high-frequency data. In multivariate price processes, we consider both…

Computational Finance · Quantitative Finance 2021-12-17 Vladimír Holý , Petra Tomanová

Imputation is a popular approach to handling censored, missing, and error-prone covariates -- all coarsened data types for which the true values are unknown. However, there are nuances to imputing these different data types based on the…

Methodology · Statistics 2025-04-29 Sarah C. Lotspeich , Ethan M. Alt

We introduce an approach to topic modelling with document-level covariates that remains tractable in the face of large text corpora. This is achieved by de-emphasizing the role of parameter estimation in an underlying probabilistic model,…

Methodology · Statistics 2025-11-05 Gabriel Phelan , David A. Campbell

We make use of the empirical process theory to approximate the adapted Hill estimator, for censored data, in terms of Gaussian processes. Then, we derive its asymptotic normality, only under the usual second-order condition of regular…

Statistics Theory · Mathematics 2015-07-07 Brahim Brahimi , Djamel Meraghni , Abdelhakim Necir

We provide some asymptotic theory for the largest eigenvalues of a sample covariance matrix of a p-dimensional time series where the dimension p = p_n converges to infinity when the sample size n increases. We give a short overview of the…

Statistics Theory · Mathematics 2016-04-27 Richard Davis , Johannes Heiny , Thomas Mikosch , Xiaolei Xie

Estimation of tail quantities, such as expected shortfall or Value at Risk, is a difficult problem. We show how the theory of nonlinear expectations, in particular the Data-robust expectation introduced in [5], can assist in the…

Statistics Theory · Mathematics 2018-02-15 Samuel N. Cohen