English

Unsupervised deep learning identifies semantic disentanglement in single inferotemporal neurons

Neurons and Cognition 2022-01-19 v1

Abstract

Deep supervised neural networks trained to classify objects have emerged as popular models of computation in the primate ventral stream. These models represent information with a high-dimensional distributed population code, implying that inferotemporal (IT) responses are also too complex to interpret at the single-neuron level. We challenge this view by modelling neural responses to faces in the macaque IT with a deep unsupervised generative model, beta-VAE. Unlike deep classifiers, beta-VAE "disentangles" sensory data into interpretable latent factors, such as gender or hair length. We found a remarkable correspondence between the generative factors discovered by the model and those coded by single IT neurons. Moreover, we were able to reconstruct face images using the signals from just a handful of cells. This suggests that the ventral visual stream may be optimising the disentangling objective, producing a neural code that is low-dimensional and semantically interpretable at the single-unit level.

Keywords

Cite

@article{arxiv.2006.14304,
  title  = {Unsupervised deep learning identifies semantic disentanglement in single inferotemporal neurons},
  author = {Irina Higgins and Le Chang and Victoria Langston and Demis Hassabis and Christopher Summerfield and Doris Tsao and Matthew Botvinick},
  journal= {arXiv preprint arXiv:2006.14304},
  year   = {2022}
}
R2 v1 2026-06-23T16:37:09.802Z