Related papers: Information Bottleneck on General Alphabets
The order of letters is not always relevant in a communication task. This paper discusses the implications of order irrelevance on source coding, presenting results in several major branches of source coding theory: lossless coding,…
Directed and undirected graphical models, also called Bayesian networks and Markov random fields, respectively, are important statistical tools in a wide variety of fields, ranging from computational biology to probabilistic artificial…
The Information Bottleneck (IB) is a conceptual method for extracting the most compact, yet informative, representation of a set of variables, with respect to the target. It generalizes the notion of minimal sufficient statistics from…
In classical information theory, the information bottleneck method (IBM) can be regarded as a method of lossy data compression which focusses on preserving meaningful (or relevant) information. As such it has recently gained a lot of…
Lower bounds for the average probability of error of estimating a hidden variable X given an observation of a correlated random variable Y, and Fano's inequality in particular, play a central role in information theory. In this paper, we…
The information bottleneck channel (or the oblivious relay channel) concerns a channel coding setting where the decoder does not directly observe the channel output. Rather, the channel output is relayed to the decoder by an oblivious relay…
We consider problems where $n$ people are communicating and a random subset of them is trying to leak information, without making it clear who are leaking the information. We introduce a measure of suspicion, and show that the amount of…
The Shannon Noiseless coding theorem (the data-compression principle) asserts that for an information source with an alphabet $\mathcal X=\{0,\ldots ,\ell -1\}$ and an asymptotic equipartition property, one can reduce the number of stored…
Linear network coding transmits data through networks by letting the intermediate nodes combine the messages they receive and forward the combinations towards their destinations. The solvability problem asks whether the demands of all the…
Information Bottleneck (IB) is widely used, but in deep learning, it is usually implemented through tractable surrogates, such as variational bounds or neural mutual information (MI) estimators, rather than directly controlling the MI…
A new notion of typicality for arbitrary probability measures on standard Borel spaces is proposed, which encompasses the classical notions of weak and strong typicality as special cases. Useful lemmas about strong typical sets, including…
As it is known, universal codes, which estimate the entropy rate consistently, exist for stationary ergodic sources over finite alphabets but not over countably infinite ones. We generalize universal coding as the problem of universal…
Invariant risk minimization (IRM) has recently emerged as a promising alternative for domain generalization. Nevertheless, the loss function is difficult to optimize for nonlinear classifiers and the original optimization objective could…
An important question for a probabilistic program is whether the probability mass of all its diverging runs is zero, that is that it terminates "almost surely". Proving that can be hard, and this paper presents a new method for doing so; it…
Let $(X_i,i\geq 1)$ be a sequence of i.i.d. random variables with values in $[0,1]$, and $f$ be a function such that $`E(f(X_1)^2)<+\infty$. We show a functional central limit theorem for the process $t\mapsto \sum_{i=1}^n f(X_i)1_{X_i\leq…
The paper establishes the central limit theorems and proposes how to perform valid inference in factor models. We consider a setting where many counties/regions/assets are observed for many time periods, and when estimation of a global…
Building upon the work of Chebyshev, Shannon and Kontoyiannis, it may be demonstrated that Chebyshev's asymptotic result: \begin{equation} \ln N \sim \sum_{p \leq N} \frac{1}{p} \cdot \ln p \end{equation} has a natural information-theoretic…
We analyze the fluctuations of incomplete $U$-statistics over a triangular array of independent random variables. We give criteria for a Central Limit Theorem (CLT, for short) to hold in the sense that we prove that an appropriately scaled…
Although deep neural networks have been immensely successful, there is no comprehensive theoretical understanding of how they work or are structured. As a result, deep networks are often seen as black boxes with unclear interpretations and…
We prove a central limit theorem for a certain class of functions on sparse rank-one inhomogeneous random graphs endowed with additional i.i.d. edge and vertex weights. Our proof of the central limit theorem uses a perturbative form of…