English

On the Difference Between the Information Bottleneck and the Deep Information Bottleneck

Machine Learning 2020-02-19 v1 Information Theory math.IT Machine Learning

Abstract

Combining the Information Bottleneck model with deep learning by replacing mutual information terms with deep neural nets has proved successful in areas ranging from generative modelling to interpreting deep neural networks. In this paper, we revisit the Deep Variational Information Bottleneck and the assumptions needed for its derivation. The two assumed properties of the data XX, YY and their latent representation TT take the form of two Markov chains TXYT-X-Y and XTYX-T-Y. Requiring both to hold during the optimisation process can be limiting for the set of potential joint distributions P(X,Y,T)P(X,Y,T). We therefore show how to circumvent this limitation by optimising a lower bound for I(T;Y)I(T;Y) for which only the latter Markov chain has to be satisfied. The actual mutual information consists of the lower bound which is optimised in DVIB and cognate models in practice and of two terms measuring how much the former requirement TXYT-X-Y is violated. Finally, we propose to interpret the family of information bottleneck models as directed graphical models and show that in this framework the original and deep information bottlenecks are special cases of a fundamental IB model.

Keywords

Cite

@article{arxiv.1912.13480,
  title  = {On the Difference Between the Information Bottleneck and the Deep Information Bottleneck},
  author = {Aleksander Wieczorek and Volker Roth},
  journal= {arXiv preprint arXiv:1912.13480},
  year   = {2020}
}
R2 v1 2026-06-23T13:00:11.222Z