English

Machine Learning Regression of stellar effective temperatures in the second $Gaia$ Data Release

Solar and Stellar Astrophysics 2019-08-14 v1 Instrumentation and Methods for Astrophysics

Abstract

This paper reports on the application of the supervised machine-learning algorithm to the stellar effective temperature regression for the second GaiaGaia data release, based on the combination of the stars in four spectroscopic surveys: Large Sky Area Multi-Object Fiber Spectroscopic Telescope, Sloan Extension for Galactic Understanding and Exploration, the Apache Point Observatory Galactic Evolution Experiment and the RAdial Velocity Extension. This combination, about four million stars, enables us to construct one of the largest training sample for the regression, and further predict reliable stellar temperatures with a root-mean-squared error of 191 K. This result is more precise than that given by GaiaGaia second data release that is based on about sixty thousands stars. After a series of data cleaning processes, the input features that feed the regressor are carefully selected from the GaiaGaia parameters, including the colors, the 3D position and the proper motion. These GaiaGaia parameters is used to predict effective temperatures for 132,739,323 valid stars in the second GaiaGaia data release. We also present a new method for blind tests and a test for external regression without additional data. The machine-learning algorithm fed with the parameters only in one catalog provides us an effective approach to maximize sample size for prediction, and this methodology has a wide application prospect in future studies of astrophysics.

Keywords

Cite

@article{arxiv.1906.09695,
  title  = {Machine Learning Regression of stellar effective temperatures in the second $Gaia$ Data Release},
  author = {Yu Bai and JiFeng Liu and ZhongRui Bai and Song Wang and DongWei Fan},
  journal= {arXiv preprint arXiv:1906.09695},
  year   = {2019}
}
R2 v1 2026-06-23T10:01:21.347Z