Related papers: Why Machine Learning Models Systematically Underes…
The difference ("mismatch") between two gravitational-wave (GW) signals is often used to estimate the signal-to-noise ratio (SNR) at which they will be distinguishable in a measurement or, alternatively, when the errors in a signal model…
A key challenge in machine learning is to explain how learning dynamics select among the many solutions that achieve identical loss values in overparameterized models - a phenomenon known as implicit bias. Controlling this bias provides a…
It has been shown that deep learning models can under certain circumstances outperform traditional statistical methods at forecasting. Furthermore, various techniques have been developed for quantifying the forecast uncertainty (prediction…
An ever-looming threat to astronomical applications of machine learning is the danger of over-fitting data, also known as the `curse of dimensionality.' This occurs when there are fewer samples than the number of independent variables. In…
Current and forthcoming cosmological data analyses share the challenge of huge datasets alongside increasingly tight requirements on the precision and accuracy of extracted cosmological parameters. The community is becoming increasingly…
In learning with noisy labels, the sample selection approach is very popular, which regards small-loss data as correctly labeled during training. However, losses are generated on-the-fly based on the model being trained with noisy labels,…
Demographic studies of cosmic populations must contend with measurement errors and selection effects. We survey some of the key ideas astronomers have developed to deal with these complications, in the context of galaxy surveys and the…
Most machine learning techniques are based upon statistical learning theory, often simplified for the sake of computing speed. This paper is focused on the uncertainty aspect of mathematical modeling in machine learning. Regression analysis…
Anomaly detection is concerned with identifying data patterns that deviate remarkably from the expected behaviour. This is an important research problem, due to its broad set of application domains, from data analysis to e-health,…
We applied machine learning to the entire data history of ESO's High Accuracy Radial Velocity Planet Searcher (HARPS) instrument. Our primary goal was to recover the physical properties of the observed objects, with a secondary emphasis on…
Weak gravitational lensing has the potential to constrain cosmological parameters to high precision. However, as shown by the Shear TEsting Programmes (STEP) and GRavitational lEnsing Accuracy Testing (GREAT) Challenges, measuring galaxy…
Increasingly large areas in cosmic shear surveys lead to a reduction of statistical errors, necessitating to control systematic errors increasingly better. One of these systematic effects was initially studied by Hartlap et al. in 2011,…
The rapid growth of large-scale radio surveys, generating over 100 petabytes of data annually, has created a pressing need for automated data analysis methods. Recent research has explored the application of machine learning techniques to…
Advances in machine learning and the increasing availability of high-dimensional data have led to the proliferation of social science research that uses the predictions of machine learning models as proxies for measures of human activity or…
Gravitational-wave (GW) observations of binary black-hole (BBH) coalescences are expected to address outstanding questions in astrophysics, cosmology, and fundamental physics. Realizing the full discovery potential of upcoming…
Nonlinearity in many systems is heavily dependent on component variation and environmental factors such as temperature. This is often overcome by keeping signals close enough to the device's operating point that it appears approximately…
The Laser Interferometer Space Antenna (LISA) will observe massive black hole binaries (MBHBs) with astoundingly high signal-to-noise ratio, leaving parameter estimation with these signals susceptible to seemingly small waveform errors. Of…
Noise in data appears to be inevitable in most real-world machine learning applications and would cause severe overfitting problems. Not only can data features contain noise, but labels are also prone to be noisy due to human input. In this…
We consider the dynamic linear regression problem, where the predictor vector may vary with time. This problem can be modeled as a linear dynamical system, with non-constant observation operator, where the parameters that need to be learned…
Given a supervised machine learning problem where the training set has been subject to a known sampling bias, how can a model be trained to fit the original dataset? We achieve this through the Bayesian inference framework by altering the…