English
Related papers

Related papers: Generalized Statistics Framework for Rate Distorti…

200 papers

A general class of Bayesian lower bounds when the underlying loss function is a Bregman divergence is demonstrated. This class can be considered as an extension of the Weinstein--Weiss family of bounds for the mean squared error and relies…

Information Theory · Computer Science 2020-06-17 Alex Dytso , Michael Fauß , H. Vincent Poor

Theoretically understanding stochastic gradient descent (SGD) in overparameterized models has led to the development of several optimization algorithms that are widely used in practice today. Recent work by~\citet{zou2021benign} provides…

Machine Learning · Computer Science 2025-06-19 Alexandru Meterez , Depen Morwani , Costin-Andrei Oncescu , Jingfeng Wu , Cengiz Pehlevan , Sham Kakade

Stochastic gradient descent (SGD) is the optimization algorithm of choice in many machine learning applications such as regularized empirical risk minimization and training deep neural networks. The classical convergence analysis of SGD is…

Optimization and Control · Mathematics 2018-07-10 Lam M. Nguyen , Phuong Ha Nguyen , Marten van Dijk , Peter Richtárik , Katya Scheinberg , Martin Takáč

Computing the rate-distortion function for continuous sources is commonly regarded as a standard continuous optimization problem. When numerically addressing this problem, a typical approach involves discretizing the source space and…

Information Theory · Computer Science 2024-05-02 Lingyi Chen , Shitong Wu , Wenyi Zhang , Huihui Wu , Hao Wu

The theory of large deviations constitutes a mathematical cornerstone in the foundations of Boltzmann-Gibbs statistical mechanics, based on the additive entropy $S_{BG}=- k_B\sum_{i=1}^W p_i \ln p_i$. Its optimization under appropriate…

Statistical Mechanics · Physics 2011-10-31 Guiomar Ruiz , Constantino Tsallis

A general expression for the distortion rate function (DRF) of cyclostationary Gaussian processes in terms of their spectral properties is derived. This expression can be seen as the result of orthogonalization over the different components…

Information Theory · Computer Science 2016-08-11 Alon Kipnis , Andrea J. Goldsmith , Yonina C. Eldar

Regularisation theory in Banach spaces, and non--norm-squared regularisation even in finite dimensions, generally relies upon Bregman divergences to replace norm convergence. This is comparable to the extension of first-order optimisation…

Optimization and Control · Mathematics 2021-03-19 Tuomo Valkonen

In this paper a new family of minimum divergence estimators based on the Bregman divergence is proposed. The popular density power divergence (DPD) class of estimators is a sub-class of Bregman divergences. We propose and study a new…

Statistics Theory · Mathematics 2020-08-18 Soumik Purkayastha , Ayanendranath Basu

We derive a refined conjecture for the variance of Gaussian primes across sectors, with a power saving error term, by applying the L-functions Ratios Conjecture. We observe a bifurcation point in the main term, consistent with the Random…

A parametric theory of statistical inference is developed for the moderate deviation probability zone. The new approach to the proofs is based on the Taylor series expansion of the logarithm of the likelihood ratio based on the Hellinger…

Statistics Theory · Mathematics 2026-04-28 Mikhail Ermakov

We introduce a temperature into the exponential function and replace the softmax output layer of neural nets by a high temperature generalization. Similarly, the logarithm in the log loss we use for training is replaced by a low temperature…

Machine Learning · Computer Science 2019-09-24 Ehsan Amid , Manfred K. Warmuth , Rohan Anil , Tomer Koren

Statistical learning theory is the foundation of machine learning, providing theoretical bounds for the risk of models learned from a (single) training set, assumed to issue from an unknown probability distribution. In actual deployment,…

Machine Learning · Computer Science 2024-10-25 Michele Caprio , Maryam Sultana , Eleni Elia , Fabio Cuzzolin

We study the common continual learning setup where an overparameterized model is sequentially fitted to a set of jointly realizable tasks. We analyze forgetting, defined as the loss on previously seen tasks, after $k$ iterations. For…

Machine Learning · Computer Science 2026-01-05 Itay Evron , Ran Levinstein , Matan Schliserman , Uri Sherman , Tomer Koren , Daniel Soudry , Nathan Srebro

Direct evaluation of the rate-distortion function has rarely been achieved when it is strictly greater than its Shannon lower bound. In this paper, we consider the rate-distortion function for the distortion measure defined by an…

Information Theory · Computer Science 2013-02-27 Kazuho Watanabe

The notion of signal sparsity has been gaining increasing interest in information theory and signal processing communities. As a consequence, a plethora of sparsity metrics has been presented in the literature. The appropriateness of these…

Information Theory · Computer Science 2016-02-08 Anastasios Maronidis , Elisavet Chatzilari , Spiros Nikolopoulos , Ioannis Kompatsiaris

We prove a general theorem to bound the total variation distance between the distribution of an integer valued random variable of interest and an appropriate discretized normal distribution. We apply the theorem to 2-runs in a sequence of…

Probability · Mathematics 2014-07-07 Xiao Fang

We propose a novel Bregman descent algorithm for minimizing a convex function that is expressed as the sum of a differentiable part (defined over an open set) and a possibly nonsmooth term. The approach, referred to as the Variable Bregman…

Machine Learning · Computer Science 2025-02-06 Ségolène Martin , Jean-Christophe Pesquet , Gabriele Steidl , Ismail Ben Ayed

Generalization error bounds for deep neural networks trained by stochastic gradient descent (SGD) are derived by combining a dynamical control of an appropriate parameter norm and the Rademacher complexity estimate based on parameter norms.…

Machine Learning · Computer Science 2023-05-30 Mingze Wang , Chao Ma

In traditional statistical learning, data points are usually assumed to be independently and identically distributed (i.i.d.) following an unknown probability distribution. This paper presents a contrasting viewpoint, perceiving data points…

Machine Learning · Computer Science 2025-08-19 Yangchen Pan , Junfeng Wen , Chenjun Xiao , Philip Torr

We develop a general optimization-theoretic framework for Bregman-Variational Learning Dynamics (BVLD), a new class of operator-based updates that unify Bayesian inference, mirror descent, and proximal learning under time-varying…

Optimization and Control · Mathematics 2025-10-24 Jinho Cha , Youngchul Kim , Jungmin Shin , Jaeyoung Cho , Seon Jin Kim , Junyeol Ryu