English
Related papers

Related papers: Information-theoretic applications of the logarith…

200 papers

We develop a unified Data Processing Inequality PAC-Bayesian framework -- abbreviated DPI-PAC-Bayesian -- for deriving generalization error bounds in the supervised learning setting. By embedding the Data Processing Inequality (DPI) into…

Information Theory · Computer Science 2025-08-26 Muhan Guan , Farhad Farokhi , Jingge Zhu

We derive tight and computable bounds on the bias of statistical estimators, or more generally of quantities of interest, when evaluated on a baseline model P rather than on the typically unknown true model Q. Our proposed method combines…

Information Theory · Computer Science 2017-07-04 Konstantinos Gourgoulias , Markos A. Katsoulakis , Luc Rey-Bellet , Jie Wang

Mutual Information (MI) is a fundamental measure of statistical dependence widely used in representation learning. While direct optimization of MI via its definition as a Kullback-Leibler divergence (KLD) is often intractable, many recent…

Machine Learning · Computer Science 2026-03-18 Reuben Dorent , Polina Golland , William Wells

The support recovery problem consists of determining a sparse subset of a set of variables that is relevant in generating a set of observations, and arises in a diverse range of settings such as compressive sensing, and subset selection in…

Information Theory · Computer Science 2016-08-31 Jonathan Scarlett , Volkan Cevher

Bounds on information combining are a fundamental tool in coding theory, in particular when analyzing polar codes and belief propagation. They usually bound the evolution of random variables with respect to their Shannon entropy. In recent…

Information Theory · Computer Science 2023-05-05 Christoph Hirche , Xinyue Guan , Marco Tomamichel

We introduce an axiomatic approach for channel divergences and channel relative entropies that is based on three information-theoretic axioms of monotonicity under superchannels (i.e. generalized data processing inequality), additivity…

Quantum Physics · Physics 2021-01-25 Gilad Gour

Since its introduction, the partial information decomposition (PID) has emerged as a powerful, information-theoretic technique useful for studying the structure of (potentially higher-order) interactions in complex systems. Despite its…

Information Theory · Computer Science 2023-12-11 Thomas F. Varley

We present a definition of the distance between probability distributions. Our definition is based on the $L_1$ norm on space of probability measures. We compare our distance with the well-known Kullback-Leibler divergence and with the…

General Relativity and Quantum Cosmology · Physics 2008-11-26 Robert J. Budzyński , Witold Kondracki , Andrzej Królak

This study aims to quantify and visualize the degradation of fidelity (information degradation) that inevitably accompanies the replication of information within the framework of information thermodynamics and to propose an…

Mathematical Physics · Physics 2025-11-20 Tatsuaki Tsuruyama

Cross-entropy loss is a common choice when it comes to multiclass classification tasks and language modeling in particular. Minimizing this loss results in language models of very good quality. We show that it is possible to fine-tune these…

Computation and Language · Computer Science 2019-01-16 Vadim Popov , Mikhail Kudinov

This manuscript introduces the idea of using Distributionally Robust Optimization (DRO) for the Counterfactual Risk Minimization (CRM) problem. Tapping into a rich existing literature, we show that DRO is a principled tool for…

Machine Learning · Statistics 2019-12-17 Louis Faury , Ugo Tanielian , Flavian Vasile , Elena Smirnova , Elvis Dohmatob

We recall some of the history of the information-theoretic approach to deriving core results in probability theory and indicate parts of the recent resurgence of interest in this area with current progress along several interesting…

Probability · Mathematics 2022-04-28 Lampros Gavalakis , Ioannis Kontoyiannis

Deep nonlinear models pose a challenge for fitting parameters due to lack of knowledge of the hidden layer and the potentially non-affine relation of the initial and observed layers. In the present work we investigate the use of information…

Optimization and Control · Mathematics 2016-12-20 Jacob S. Hunter , Nathan O. Hodas

A loss function measures the discrepancy between the true values and their estimated fits, for a given instance of data. In classification problems, a loss function is said to be proper if a minimizer of the expected loss is the true…

Information Theory · Computer Science 2020-01-03 Amichai Painsky , Gregory W. Wornell

We consider the problem of decision-making with side information and unbounded loss functions. Inspired by probably approximately correct learning model, we use a slightly different model that incorporates the notion of side information in…

Machine Learning · Computer Science 2007-07-13 Majid Fozunbal , Ton Kalker

We present a unified technique for sequential estimation of convex divergences between distributions, including integral probability metrics like the kernel maximum mean discrepancy, $\varphi$-divergences like the Kullback-Leibler…

Statistics Theory · Mathematics 2023-03-14 Tudor Manole , Aaditya Ramdas

We leverage the Gibbs inequality and its natural generalization to R\'enyi entropies to derive closed-form parametric expressions of the optimal lower bounds of $\rho$th-order guessing entropy (guessing moment) of a secret taking values on…

Information Theory · Computer Science 2024-01-31 Julien Béguinot , Olivier Rioul

Rare events, and more general risk-sensitive quantities-of-interest (QoIs), are significantly impacted by uncertainty in the tail behavior of a distribution. Uncertainty in the tail can take many different forms, each of which leads to a…

Probability · Mathematics 2019-11-22 Jeremiah Birrell , Paul Dupuis , Markos A. Katsoulakis , Luc Rey-Bellet , Jie Wang

Bayesian model averaging is a practical method for dealing with uncertainty due to model specification. Use of this technique requires the estimation of model probability weights. In this work, we revisit the derivation of estimators for…

Methodology · Statistics 2024-02-05 Ethan T. Neil , Jacob W. Sitison

We study the fundamental and timely problem of learning long sequences in autoregressive modeling and next-token prediction under model misspecification, measured by the joint Kullback--Leibler (KL) divergence. Our goal is to characterize…

Machine Learning · Computer Science 2026-05-13 Yunbei Xu , Yuzhe Yuan , Ruohan Zhan