English
Related papers

Related papers: Generalized Pearson correlation squares for captur…

200 papers

Measuring and quantifying dependencies between random variables (RV's) can give critical insights into a data-set. Typical questions are: `Do underlying relationships exist?', `Are some variables redundant?', and `Is some target variable…

Machine Learning · Statistics 2022-03-24 Guus Berkelmans , Joris Pries , Sandjai Bhulai , Rob van der Mei

We describe a general framework -- compressive statistical learning -- for resource-efficient large-scale learning: the training collection is compressed in one pass into a low-dimensional sketch (a vector of random empirical generalized…

Machine Learning · Statistics 2021-06-23 Rémi Gribonval , Gilles Blanchard , Nicolas Keriven , Yann Traonmilin

The extension of bivariate measures of dependence to non-Euclidean spaces is a challenging problem. The non-linear nature of these spaces makes the generalisation of classical measures of linear dependence (such as the covariance) not…

Statistics Theory · Mathematics 2024-10-10 Meshal Abuqrais , Davide Pigoli

Standard Gaussian graphical models (GGMs) implicitly assume that the conditional independence among variables is common to all observations in the sample. However, in practice, observations are usually collected form heterogeneous…

Methodology · Statistics 2010-01-26 Abel Rodriguez , Alex Lenkoski , Adrian Dobra

In this work, we present a new approach for constructing models for correlation matrices with a user-defined graphical structure. The graphical structure makes correlation matrices interpretable and avoids the quadratic increase of…

Clustering task of mixed data is a challenging problem. In a probabilistic framework, the main difficulty is due to a shortage of conventional distributions for such data. In this paper, we propose to achieve the mixed data clustering with…

Methodology · Statistics 2015-10-01 Matthieu Marbac , Christophe Biernacki , Vincent Vandewalle

Covariate-adaptive randomization (CAR) procedures are frequently used in comparative studies to increase the covariate balance across treatment groups. However, because randomization inevitably uses the covariate information when forming…

Statistics Theory · Mathematics 2022-07-08 Wei Ma , Yichen Qin , Yang Li , Feifang Hu

This paper introduces a method for studying the correlation structure of a range of responses modelled by a multivariate generalised linear mixed model (MGLMM). The methodology requires the existence of clusters of observations and that…

Methodology · Statistics 2021-08-02 Jeanett S. Pelck , Rodrigo Labouriau

In this paper we give a completely new approach to the problem of covariate selection in linear regression. A covariate or a set of covariates is included only if it is better in the sense of least squares than the same number of Gaussian…

Methodology · Statistics 2022-02-25 Laurie Davies , Lutz Dümbgen

This paper reviews generalized Pareto copulas (GPC), which turn out to be a key to multivariate extreme value theory. Any GPC can be represented in an easy analytic way using a particular type of norm on $\mathbb{R}^d$, called $D$-norm. The…

Statistics Theory · Mathematics 2018-11-26 Michael Falk , Simone Padoan , Florian Wisheckel

We present a general framework, the coupled compound Poisson factorization (CCPF), to capture the missing-data mechanism in extremely sparse data sets by coupling a hierarchical Poisson factorization with an arbitrary data-generating model.…

Machine Learning · Computer Science 2017-01-10 Mehmet E. Basbug , Barbara E. Engelhardt

Estimating causal effects in quasi-experiments with spatio-temporal panel data often requires adjusting for unmeasured confounding that varies across space and time. Gaussian Processes (GPs) offer a flexible, nonparametric modeling approach…

Methodology · Statistics 2025-07-08 Sofia L. Vega , Rachel C. Nethery

In this paper we give a completely new approach to the problem of covariate selection in linear regression. A covariate or a set of covariates is included only if it is better in the sense of least squares than the same number of Gaussian…

Methodology · Statistics 2023-02-09 Laurie Davies , Lutz Dümbgen

Prediction invariance of causal models under heterogeneous settings has been exploited by a number of recent methods for causal discovery, typically focussing on recovering the causal parents of a target variable of interest. Existing…

Methodology · Statistics 2026-03-10 Alice Polinelli , Veronica Vinciotti , Ernst C. Wit

Clustering mixtures of Gaussian distributions is a fundamental and challenging problem that is ubiquitous in various high-dimensional data processing tasks. While state-of-the-art work on learning Gaussian mixture models has focused…

Machine Learning · Computer Science 2018-03-05 Dan Kushnir , Shirin Jalali , Iraj Saniee

In this article, we develop a distributed variable screening method for generalized linear models. This method is designed to handle situations where both the sample size and the number of covariates are large. Specifically, the proposed…

Methodology · Statistics 2024-05-09 Tianbo Diao , Lianqiang Qu , Bo Li , Liuquan Sun

Many application domains such as ecology or genomics have to deal with multivariate non Gaussian observations. A typical example is the joint observation of the respective abundances of a set of species in a series of sites, aiming to…

Methodology · Statistics 2018-05-01 Julien Chiquet , Mahendra Mariadassou , Stéphane Robin

A generalization of Gy's theory for the variance of the fundamental sampling error is reviewed. Practical situations where the generalized model potentially leads to more accurate variance estimates are identified as: clustering of…

Applications · Statistics 2009-11-10 Bastiaan Geelhoed

Graphical models are an important tool in exploring relationships between variables in complex, multivariate data. Methods for learning such graphical models are well developed in the case where all variables are either continuous or…

Machine Learning · Statistics 2024-02-15 Konstantin Göbler , Anne Miloschewski , Mathias Drton , Sach Mukherjee

We investigate a generalized empirical likelihood approach in a two-group setting where the constraints on parameters have a form of U-statistics. In this situation, the summands that consist of the constraints for the empirical likelihood…

Methodology · Statistics 2015-05-04 Jihnhee Yu , Luge Yang , Albert Vexler , Alan D. Hutson