Related papers: Overfitting and correlations in model fitting with…
Weighted correlation functions are an increasingly important tool for understanding how galaxy properties depend on their separation from each other. We use a mock galaxy sample drawn from the Millenium simulation, assigning weights using a…
We obtain general, exact formulas for the overlaps between the eigenvectors of large correlated random matrices, with additive or multiplicative noise. These results have potential applications in many different contexts, from quantum…
Data analysis in cosmology requires reliable covariance matrices. Covariance matrices derived from numerical simulations often require a very large number of realizations to be accurate. When a theoretical model for the covariance matrix…
Using a time series model to mimic an observed time series has a long history. However, with regard to this objective, conventional estimation methods for discrete-time dynamical models are frequently found to be wanting. In fact, they are…
Observables in particle physics and specifically in lattice QCD calculations are often extracted from fits. Standard $\chi^2$ tests require a reliable determination of the covariance matrix and its inverse from correlated and…
Motivated by the importance ascribed to correlations in random matrices used to model phenomena in various scientific disciplines, we report how algebraic correlations between matrix elements affect the eigenvalue statistics and spectral…
Overfitting is a well-known issue in machine learning that occurs when a model struggles to generalize its predictions to new, unseen data beyond the scope of its training set. Traditional techniques to mitigate overfitting include early…
A common problem in health research is that we have a large database with many variables measured on a large number of individuals. We are interested in measuring additional variables on a subsample; these measurements may be newly…
Recent work has focused on the problem of conducting linear regression when the number of covariates is very large, potentially greater than the sample size. To facilitate this, one useful tool is to assume that the model can be well…
In modern applications, statisticians are faced with integrating heterogeneous data modalities relevant for an inference, prediction, or decision problem. In such circumstances, it is convenient to use a graphical model to represent the…
Measurement error arises through a variety of mechanisms. A rich literature exists on the bias introduced by covariate measurement error and on methods of analysis to address this bias. By comparison, less attention has been given to errors…
Deep, overparameterized regression models are notorious for their tendency to overfit. This problem is exacerbated in heteroskedastic models, which predict both mean and residual noise for each data point. At one extreme, these models fit…
Soil carbon accounting and prediction play a key role in building decision support systems for land managers selling carbon credits, in the spirit of the Paris and Kyoto protocol agreements. Land managers typically rely on computationally…
This work considers the problem of binary classification: given training data $x_1, \dots, x_n$ from a certain population, together with associated labels $y_1,\dots, y_n \in \left\{0,1 \right\}$, determine the best label for an element $x$…
The modern definition of optical coherence highlights a frequency dependent function based on a matrix of spectra and cross-spectra. Due to general properties of matrices, such a function is invariant in changes of basis. In this article,…
A superposition of a matrix ensemble refers to the ensemble constructed from two independent copies of the original, while a decimation refers to the formation of a new ensemble by observing only every second eigenvalue. In the cases of the…
Uncoupled regression is the problem to learn a model from unlabeled data and the set of target values while the correspondence between them is unknown. Such a situation arises in predicting anonymized targets that involve sensitive…
Regularized regression models are well studied and, under appropriate conditions, offer fast and statistically interpretable results. However, large data in many applications are heterogeneous in the sense of harboring distributional…
Machine learning models that are overfitted/overtrained are more vulnerable to knowledge leakage, which poses a risk to privacy. Suppose we download or receive a model from a third-party collaborator without knowing its training accuracy.…
We consider the limitations of two techniques for detecting nonlinearity in time series. The first technique compares the original time series to an ensemble of surrogate time series that are constructed to mimic the linear properties of…