English

Natural representation of composite data with replicated autoencoders

Quantitative Methods 2019-10-04 v1 Machine Learning Machine Learning

Abstract

Generative processes in biology and other fields often produce data that can be regarded as resulting from a composition of basic features. Here we present an unsupervised method based on autoencoders for inferring these basic features of data. The main novelty in our approach is that the training is based on the optimization of the `local entropy' rather than the standard loss, resulting in a more robust inference, and enhancing the performance on this type of data considerably. Algorithmically, this is realized by training an interacting system of replicated autoencoders. We apply this method to synthetic and protein sequence data, and show that it is able to infer a hidden representation that correlates well with the underlying generative process, without requiring any prior knowledge.

Keywords

Cite

@article{arxiv.1909.13327,
  title  = {Natural representation of composite data with replicated autoencoders},
  author = {Matteo Negri and Davide Bergamini and Carlo Baldassi and Riccardo Zecchina and Christoph Feinauer},
  journal= {arXiv preprint arXiv:1909.13327},
  year   = {2019}
}

Comments

11 pages, 4 figures

R2 v1 2026-06-23T11:29:30.732Z