English

Re-envisioning Euclid Galaxy Morphology: Identifying and Interpreting Features with Sparse Autoencoders

Instrumentation and Methods for Astrophysics 2025-11-13 v2 Machine Learning

Abstract

Sparse Autoencoders (SAEs) can efficiently identify candidate monosemantic features from pretrained neural networks for galaxy morphology. We demonstrate this on Euclid Q1 images using both supervised (Zoobot) and new self-supervised (MAE) models. Our publicly released MAE achieves superhuman image reconstruction performance. While a Principal Component Analysis (PCA) on the supervised model primarily identifies features already aligned with the Galaxy Zoo decision tree, SAEs can identify interpretable features outside of this framework. SAE features also show stronger alignment than PCA with Galaxy Zoo labels. Although challenges in interpretability remain, SAEs provide a powerful engine for discovering astrophysical phenomena beyond the confines of human-defined classification.

Cite

@article{arxiv.2510.23749,
  title  = {Re-envisioning Euclid Galaxy Morphology: Identifying and Interpreting Features with Sparse Autoencoders},
  author = {John F. Wu and Michael Walmsley},
  journal= {arXiv preprint arXiv:2510.23749},
  year   = {2025}
}

Comments

Authors contributed equally to this work. Accepted to NeurIPS Machine Learning and the Physical Sciences Workshop. See trained model at https://huggingface.co/mwalmsley/euclid-rr2-mae, HuggingFace demo at https://huggingface.co/spaces/mwalmsley/euclid_masked_autoencoder, and code at https://github.com/jwuphysics/euclid-galaxy-morphology-saes

R2 v1 2026-07-01T07:08:23.632Z