English
Related papers

Related papers: The Sample Complexity of Lossless Data Compression

200 papers

We establish finite sample bounds for the error of standard and waste-free SMC samplers. Our results cover estimates of both expectations and normalising constants of the target distributions. We consider first an arbitrary sequence of…

Computation · Statistics 2026-04-15 Yvann Le Fay , Nicolas Chopin , Matti Vihola

Extracting noisy or incorrectly labeled samples from a labeled dataset with hard/difficult samples is an important yet under-explored topic. Two general and often independent lines of work exist, one focuses on addressing noisy labels, and…

Machine Learning · Computer Science 2023-07-21 Mahsa Forouzesh , Patrick Thiran

We present herein a scheme by which to accurately evaluate the error exponents of a lossy data compression problem, which characterize average probabilities over a code ensemble of compression failure and success above or below a critical…

Statistical Mechanics · Physics 2007-05-23 Tadaaki Hosaka , Yoshiyuki Kabashima

Many current applications in data science need rich model classes to adequately represent the statistics that may be driving the observations. But rich model classes may be too complex to admit estimators that converge to the truth with…

Information Theory · Computer Science 2022-05-04 N. Santhanam , V. Anantharam , W. Szpankowski

A lossy compression algorithm for binary redundant memoryless sources is presented. The proposed scheme is based on sparse graph codes. By introducing a nonlinear function, redundant memoryless sequences can be compressed. We propose a…

Information Theory · Computer Science 2011-08-19 Kazushi Mimura

The paper presents exponentially-strong converses for source-coding, channel coding, and hypothesis testing problems. More specifically, it presents alternative proofs for the well-known exponentially-strong converse bounds for almost…

Information Theory · Computer Science 2023-01-18 Mustapha Hamad , Michele Wigger , Mireille Sarkiss

Traditional statistical analysis requires that the analysis process and data are independent. By contrast, the new field of adaptive data analysis hopes to understand and provide algorithms and accuracy guarantees for research as it is…

Machine Learning · Computer Science 2017-03-22 Sam Elder

Constrained decoding enables Language Models (LMs) to produce samples that provably satisfy hard constraints. However, existing constrained-decoding approaches often distort the underlying model distribution, a limitation that is especially…

Artificial Intelligence · Computer Science 2025-06-09 Emmanuel Anaya Gonzalez , Sairam Vaidya , Kanghee Park , Ruyi Ji , Taylor Berg-Kirkpatrick , Loris D'Antoni

Recent advances have significantly improved our understanding of the sample complexity of learning in average-reward Markov decision processes (AMDPs) under the generative model. However, much less is known about the constrained…

Machine Learning · Computer Science 2025-09-23 Yukuan Wei , Xudong Li , Lin F. Yang

This paper describes universal lossless coding strategies for compressing sources on countably infinite alphabets. Classes of memoryless sources defined by an envelope condition on the marginal distribution provide benchmarks for coding…

Statistics Theory · Mathematics 2015-01-05 Stéphane Boucheron , Aurélien Garivier , Elisabeth Gassiat

We design and mathematically analyze sampling-based algorithms for regularized loss minimization problems that are implementable in popular computational models for large data, in which the access to the data is restricted in some way. Our…

Machine Learning · Computer Science 2019-06-04 Ryan R. Curtin , Sungjin Im , Ben Moseley , Kirk Pruhs , Alireza Samadian

We review several statistical complexity measures proposed over the last decade and a half as general indicators of structure or correlation. Recently, Lopez-Ruiz, Mancini, and Calbet [Phys. Lett. A 209 (1995) 321] introduced another…

Statistical Mechanics · Physics 2008-02-03 David P. Feldman , James P. Crutchfield

We consider the problem of approximating a function in a general nonlinear subset of $L^2$, when only a weighted Monte Carlo estimate of the $L^2$-norm can be computed. Of particular interest in this setting is the concept of sample…

Numerical Analysis · Mathematics 2023-01-24 Philipp Trunschke

We consider the problem of estimating the frequency components of a mixture of s complex sinusoids from a random subset of n regularly spaced samples. Unlike previous work in compressed sensing, the frequencies are not assumed to lie on a…

Information Theory · Computer Science 2013-07-12 Gongguo Tang , Badri Narayan Bhaskar , Parikshit Shah , Benjamin Recht

The problem of channel coding with the erasure option is revisited for discrete memoryless channels. The interplay between the code rate, the undetected and total error probabilities is characterized. Using the information spectrum method,…

Information Theory · Computer Science 2015-10-22 Masahito Hayashi , Vincent Y. F. Tan

We study the sample complexity of the Sign-Perturbed Sums (SPS) method, which constructs exact, non-asymptotic confidence regions for the true system parameters under mild statistical assumptions, such as independent and symmetric noise…

Machine Learning · Statistics 2024-09-04 Szabolcs Szentpéteri , Balázs Csanád Csáji

It is shown that an i.i.d. binary source sequence $X_1, \ldots, X_n$ can be losslessly compressed at any rate above entropy such that the individual decoding of any $X_i$ reveals \emph{no} information about the other bits $\{X_j : j \neq…

Information Theory · Computer Science 2025-11-19 Venkat Chandar , Aslan Tchamkerten , Shashank Vatedka

We study the redundancy of universally compressing strings $X_1,\dots, X_n$ generated by a binary Markov source $p$ without any bound on the memory. To better understand the connection between compression and estimation in the Markov…

Information Theory · Computer Science 2018-06-21 Changlong Wu , Maryam Hosseini , Narayana Santhanam

This paper reports, by way of introduction, on the advances made by our group and the broader signal processing community on the concept of sample abundance; a phenomenon that naturally arises in one-bit and few-bit signal processing…

Information Theory · Computer Science 2025-07-28 Arian Eamaz , Farhang Yeganegi , Mojtaba Soltanalian

We consider a Shannon cipher system for memoryless sources, in which distortion is allowed at the legitimate decoder. The source is compressed using a rate distortion code secured by a shared key, which satisfies a constraint on the…

Information Theory · Computer Science 2016-11-17 Nir Weinberger , Neri Merhav