Related papers: Overfitting and correlations in model fitting with…
Copulas are essential tools in statistics and probability theory, enabling the study of the dependence structure between random variables independently of their marginal distributions. Among the various types of copulas, Ratio-Type Copulas…
We consider the problem of predicting several response variables using the same set of explanatory variables. This setting naturally induces a group structure over the coefficient matrix, in which every explanatory variable corresponds to a…
This study is about inducing classifiers using data that is imbalanced, with a minority class being under-represented in relation to the majority classes. The first section of this research focuses on the main characteristics of data that…
To what extent can agents with misspecified subjective models predict false correlations? We study an "analyst" who utilizes models that take the form of a recursive system of linear regression equations. The analyst fits each equation to…
Indistinguishable quantum emitters confined to length scales smaller than the wavelength of the light become superradiant. Compared to uncorrelated and distinguishable emitters, superradiance results in qualitative modifications of optical…
Data analysis based on information from several sources is common in economic and biomedical studies. This setting is often referred to as the data fusion problem, which differs from traditional missing data problems since no complete data…
We propose new methods for multivariate linear regression when the regression coefficient matrix is sparse and the error covariance matrix is dense. We assume that the error covariance matrix has equicorrelation across the response…
High dimensional error covariance matrices and their inverses are used to weight the contribution of observation and background information in data assimilation procedures. As observation error covariance matrices are often obtained by…
We introduce the coverage correlation coefficient, a novel nonparametric measure of statistical association designed to quantifies the extent to which two random variables have a joint distribution concentrated on a singular subset with…
Relational models generalize log-linear models to arbitrary discrete sample spaces by specifying effects associated with any subsets of their cells. A relational model may include an overall effect, pertaining to every cell after a…
Using the LRT statistic, a model R^2 is proposed for the generalized linear mixed model for assessing the association between the correlated outcomes and fixed effects. The R^2 compares the full model to a null model with all fixed effects…
Aims. To describe the theory of surface layer independent model fitting by phase matching and to apply this to the stars HD49933 observed by CoRoT, and HD177153 (aka Perky), observed by Kepler Methods. We use theoretical analysis, phase…
In matched observational studies with continuous treatments, individuals with different treatment doses but the same or similar covariate values are paired for causal inference. While inexact covariate matching (i.e., covariate imbalance…
The maximal correlation coefficient is a well-established generalization of the Pearson correlation coefficient for measuring non-linear dependence between random variables. It is appealing from a theoretical standpoint, satisfying…
We compute exactly the overlap between the eigenvectors of two large empirical covariance matrices computed over intersecting time intervals, generalizing the results obtained previously for non-intersecting intervals. Our method relies on…
We investigate the impact of limited data on training pairwise energy-based models for inverse problems aimed at identifying interaction networks. Utilizing the Gaussian model as testbed, we dissect training trajectories across the…
Missing covariate data commonly occur in epidemiological and clinical research, and are often dealt with using multiple imputation (MI). Imputation of partially observed covariates is complicated if the substantive model is non-linear (e.g.…
Mixtures of Linear Regressions (MLR) is an important mixture model with many applications. In this model, each observation is generated from one of the several unknown linear regression components, where the identity of the generated…
The Cox proportional hazards model is ubiquitous in the analysis of time-to-event data. However, when the data dimension p is comparable to the sample size $N$, maximum likelihood estimates for its regression parameters are known to be…
Correlation matrices are widely used to analyze the interdependence of variables in various real-world scenarios. Often, a perturbation in a few variables leads to mild differences in many correlation coefficients associated with these…