English

Combating the Instability of Mutual Information-based Losses via Regularization

Machine Learning 2022-06-22 v4 Information Theory math.IT Machine Learning

Abstract

Notable progress has been made in numerous fields of machine learning based on neural network-driven mutual information (MI) bounds. However, utilizing the conventional MI-based losses is often challenging due to their practical and mathematical limitations. In this work, we first identify the symptoms behind their instability: (1) the neural network not converging even after the loss seemed to converge, and (2) saturating neural network outputs causing the loss to diverge. We mitigate both issues by adding a novel regularization term to the existing losses. We theoretically and experimentally demonstrate that added regularization stabilizes training. Finally, we present a novel benchmark that evaluates MI-based losses on both the MI estimation power and its capability on the downstream tasks, closely following the pre-existing supervised and contrastive learning settings. We evaluate six different MI-based losses and their regularized counterparts on multiple benchmarks to show that our approach is simple yet effective.

Keywords

Cite

@article{arxiv.2011.07932,
  title  = {Combating the Instability of Mutual Information-based Losses via Regularization},
  author = {Kwanghee Choi and Siyeong Lee},
  journal= {arXiv preprint arXiv:2011.07932},
  year   = {2022}
}

Comments

Kwanghee Choi and Siyeong Lee contributed equally to this paper. Accepted for the 38th Conference on Uncertainty in Artificial Intelligence (UAI 2022)

R2 v1 2026-06-23T20:16:57.116Z