English

Properties of Minimizing Entropy

Machine Learning 2021-12-07 v1

Abstract

Compact data representations are one approach for improving generalization of learned functions. We explicitly illustrate the relationship between entropy and cardinality, both measures of compactness, including how gradient descent on the former reduces the latter. Whereas entropy is distribution sensitive, cardinality is not. We propose a third compactness measure that is a compromise between the two: expected cardinality, or the expected number of unique states in any finite number of draws, which is more meaningful than standard cardinality as it discounts states with negligible probability mass. We show that minimizing entropy also minimizes expected cardinality.

Keywords

Cite

@article{arxiv.2112.03143,
  title  = {Properties of Minimizing Entropy},
  author = {Xu Ji and Lena Nehale-Ezzine and Maksym Korablyov},
  journal= {arXiv preprint arXiv:2112.03143},
  year   = {2021}
}
R2 v1 2026-06-24T08:06:11.500Z