Related papers: Missing Data Imputation for Galaxy Redshift Estima…
Given multiband photometric data from the SDSS DR6, we estimate galaxy redshifts. We employ a Random Forest trained on color features and spectroscopic redshifts from 80,000 randomly chosen primary galaxies yielding a mapping from color to…
A common problem faced by statistical institutes is that data may be missing from collected data sets. The typical way to overcome this problem is to impute the missing data. The problem of imputing missing data is complicated by the fact…
Missing data is a common problem in real-world settings and particularly relevant in healthcare applications where researchers use Electronic Health Records (EHR) and results of observational studies to apply analytics methods. This issue…
Flagship near-future surveys targeting $10^8-10^9$ galaxies across cosmic time will soon reveal the processes of galaxy assembly in unprecedented resolution. This creates an immediate computational challenge on effective analyses of the…
This research deals with the estimation and imputation of missing data in longitudinal models with a Poisson response variable inflated with zeros. A methodology is proposed that is based on the use of maximum likelihood, assuming that data…
The number density of galaxy clusters across mass and redshift has been established as a powerful cosmological probe. Cosmological analyses with galaxy clusters traditionally employ scaling relations. However, many challenges arise from…
Incomplete observability of data generates an identification problem. There is no panacea for missing data. What one can learn about a population parameter depends on the assumptions one finds credible to maintain. The credibility of…
Accurate and automated galaxy redshift determination is essential for maximizing the scientific return of spectroscopic surveys. In this paper, we propose a data-driven method to address this challenge. The method first learns a rest-frame…
This paper tackles the problem of missing data imputation for noisy and non-Gaussian data. A classical imputation method, the Expectation Maximization (EM) algorithm for Gaussian mixture models, has shown interesting properties when…
The NASA/IPAC Extragalactic Database (NED) is an impressive tool for finding near-exhaustive information on millions of astrophysical objects. Here, we outline a small systematic error that occurs in NED because a low-redshift approximation…
The problem of missing data, usually absent incurated and competition-standard datasets, is an unfortunate reality for most machine learning models used in industry applications. Recent work has focused on understanding the nature and the…
Photometric redshift estimation is becoming an increasingly important technique, although the currently existing methods present several shortcomings which hinder their application. Here it is shown that most of those drawbacks are…
Widely used methods for analyzing missing data can be biased in small samples. To understand these biases, we evaluate in detail the situation where a small univariate normal sample, with values missing at random, is analyzed using either…
Data analysis usually suffers from the Missing Not At Random (MNAR) problem, where the cause of the value missing is not fully observed. Compared to the naive Missing Completely At Random (MCAR) problem, it is more in line with the…
We explore the accuracy of the clustering-based redshift inference within the MICE2 simulation. This method uses the spatial clustering of galaxies between a spectroscopic reference sample and an unknown sample. The goal of this study is to…
Imputing missing values is an important preprocessing step in data analysis, but the literature offers little guidance on how to choose between different imputation models. This letter suggests adopting the imputation model that generates a…
Missing data are a common problem for both the construction and implementation of a prediction algorithm. Pattern mixture kernel submodels (PMKS) - a series of submodels for every missing data pattern that are fit using only data from that…
The determination of galaxy merger fraction of field galaxies using automatic morphological indices and photometric redshifts is affected by several biases if observational errors are not properly treated. Here, we correct these biases…
Cosmology inference of galaxy clustering at the field level with the EFT likelihood in principle allows for extracting all non-Gaussian information from quasi-linear scales, while robustly marginalizing over any astrophysical uncertainties.…
Many physical properties of galaxies correlate with one another, and these correlations are often used to constrain galaxy formation models. Such correlations include the color-magnitude relation, the luminosity-size relation, the…