Building robust medical machine learning systems requires pretraining strategies that exploit the intrinsic structure present in clinical data. We introduce Multiview Masked Autoencoder (MVMAE), a self-supervised framework that leverages the natural multi-view organization of radiology studies to learn view-invariant and disease-relevant representations. MVMAE combines masked image reconstruction with cross-view alignment, transforming clinical redundancy across projections into a powerful self-supervisory signal. We further extend this approach with MVMAE-V2T, which incorporates radiology reports as an auxiliary text-based learning signal to enhance semantic grounding while preserving fully vision-based inference. Evaluated on a downstream disease classification task on three large-scale public datasets, MIMIC-CXR, CheXpert, and PadChest, MVMAE consistently outperforms supervised and vision-language baselines. Furthermore, MVMAE-V2T provides additional gains, particularly in low-label regimes where structured textual supervision is most beneficial. Together, these results establish the importance of structural and textual supervision as complementary paths toward scalable, clinically grounded medical foundation models.
@article{arxiv.2511.22294,
title = {Structure is Supervision: Multiview Masked Autoencoders for Radiology},
author = {Sonia Laguna and Andrea Agostini and Alain Ryser and Samuel Ruiperez-Campillo and Irene Cannistraci and Moritz Vandenhirtz and Stephan Mandt and Nicolas Deperrois and Farhad Nooralahzadeh and Michael Krauthammer and Thomas M. Sutter and Julia E. Vogt},
journal= {arXiv preprint arXiv:2511.22294},
year = {2026}
}