English

Cycle-StarNet: Bridging the gap between theory and data by leveraging large datasets

Solar and Stellar Astrophysics 2021-01-22 v3 Astrophysics of Galaxies Instrumentation and Methods for Astrophysics Data Analysis, Statistics and Probability Machine Learning

Abstract

The advancements in stellar spectroscopy data acquisition have made it necessary to accomplish similar improvements in efficient data analysis techniques. Current automated methods for analyzing spectra are either (a) data-driven, which requires prior knowledge of stellar parameters and elemental abundances, or (b) based on theoretical synthetic models that are susceptible to the gap between theory and practice. In this study, we present a hybrid generative domain adaptation method that turns simulated stellar spectra into realistic spectra by applying unsupervised learning to large spectroscopic surveys. We apply our technique to the APOGEE H-band spectra at R=22,500 and the Kurucz synthetic models. As a proof of concept, two case studies are presented. The first of which is the calibration of synthetic data to become consistent with observations. To accomplish this, synthetic models are morphed into spectra that resemble observations, thereby reducing the gap between theory and observations. Fitting the observed spectra shows an improved average reduced χR2\chi_R^2 from 1.97 to 1.22, along with a reduced mean residual from 0.16 to -0.01 in normalized flux. The second case study is the identification of the elemental source of missing spectral lines in the synthetic modelling. A mock dataset is used to show that absorption lines can be recovered when they are absent in one of the domains. This method can be applied to other fields, which use large data sets and are currently limited by modelling accuracy. The code used in this study is made publicly available on github.

Keywords

Cite

@article{arxiv.2007.03109,
  title  = {Cycle-StarNet: Bridging the gap between theory and data by leveraging large datasets},
  author = {Teaghan O'Briain and Yuan-Sen Ting and Sébastien Fabbro and Kwang M. Yi and Kim Venn and Spencer Bialek},
  journal= {arXiv preprint arXiv:2007.03109},
  year   = {2021}
}

Comments

23 pages, 15 figures, 2 tables, accepted for publication on Nov 12, 2020, Nov 12. A companion 4-page preview is accepted to the ICML 2020 Machine Learning Interpretability for Scientific Discovery workshop. The code used in this study is made publicly available on github: https://github.com/teaghan/Cycle_SN

R2 v1 2026-06-23T16:54:06.964Z