English
Related papers

Related papers: A fast algorithm for computing distance correlatio…

200 papers

This paper addresses the task of estimating a covariance matrix under a patternless sparsity assumption. In contrast to existing approaches based on thresholding or shrinkage penalties, we propose a likelihood-based method that regularizes…

Methodology · Statistics 2021-09-13 Jason Xu , Kenneth Lange

Distance correlation is a recent extension of Pearson's correlation, that characterises general statistical independence between Euclidean-space-valued random variables, not only linear relations. This review delves into how and when…

Statistics Theory · Mathematics 2020-09-30 Fernando Castro-Prado , Wenceslao González-Manteiga

Distance queries are a basic tool in data analysis. They are used for detection and localization of change for the purpose of anomaly detection, monitoring, or planning. Distance queries are particularly useful when data sets such as…

Data Structures and Algorithms · Computer Science 2015-03-20 Edith Cohen

Non-parametric two-sample tests based on energy distance or maximum mean discrepancy are widely used statistical tests for comparing multivariate data from two populations. While these tests enjoy desirable statistical properties, their…

Computation · Statistics 2024-06-11 Elias Chaibub Neto

The problem of testing changes in covariance has received increasing attention in recent years, especially in the context of high-dimensional testing. A number of approaches have been proposed, all limited to the two-sample problem and…

Methodology · Statistics 2016-09-06 Yi-Hui Zhou

The total variation distance is a metric of central importance in statistics and probability theory. However, somewhat surprisingly, questions about computing it algorithmically appear not to have been systematically studied until very…

Data Structures and Algorithms · Computer Science 2025-03-17 Arnab Bhattacharyya , Weiming Feng , Piyush Srivastava

This paper investigates the utilization of maximum and average distance correlations for multivariate independence testing. We characterize their consistency properties in high-dimensional settings with respect to the number of marginally…

Machine Learning · Statistics 2025-06-11 Cencheng Shen , Yuexiao Dong

Our research proposes a novel method for reducing the dimensionality of functional data, specifically for the case where the response is a scalar and the predictor is a random function. Our method utilizes distance covariance, and has…

Statistics Theory · Mathematics 2023-09-26 Xing Yang , Jianjun Xu

Consider a set $P$ of $n$ points picked uniformly and independently from $[0,1]^d$ for a constant dimension $d$ -- such a point set is extremely well behaved in many aspects. For example, for a fixed $r \in [0,1]$, we prove a new…

Computational Geometry · Computer Science 2023-11-01 Sariel Har-Peled , Elfarouk Harb

Pearson's Chi-squared test, though widely used for detecting association between categorical variables, exhibits low statistical power in large sparse contingency tables. To address this limitation, two novel permutation tests have been…

Methodology · Statistics 2024-03-27 Qingyang Zhang

We analyze how the choice of the sampling weight affects the efficiency of the Monte Carlo evaluation of classical time autocorrelation functions. Assuming uncorrelated sampling or sampling with constant correlation length, we propose a…

Chemical Physics · Physics 2015-06-12 Tomas Zimmermann , Jiri Vanicek

We study the problem of detecting outlier pairs of strongly correlated variables among a collection of $n$ variables with otherwise weak pairwise correlations. After normalization, this task amounts to the geometric task where we are given…

Data Structures and Algorithms · Computer Science 2018-01-08 Matti Karppa , Petteri Kaski , Jukka Kohonen

In data science, it is often required to estimate dependencies between different data sources. These dependencies are typically calculated using Pearson's correlation, distance correlation, and/or mutual information. However, none of these…

Statistics Theory · Mathematics 2015-06-03 Rahul Agarwal , Pierre Sacre , Sridevi V. Sarma

We provide an algorithm for properly learning mixtures of two single-dimensional Gaussians without any separability assumptions. Given $\tilde{O}(1/\varepsilon^2)$ samples from an unknown mixture, our algorithm outputs a mixture that is…

Data Structures and Algorithms · Computer Science 2014-05-20 Constantinos Daskalakis , Gautam Kamath

Covariance matrix estimation concerns the problem of estimating the covariance matrix from a collection of samples, which is of extreme importance in many applications. Classical results have shown that $O(n)$ samples are sufficient to…

Information Theory · Computer Science 2019-03-19 Wei Cui , Xu Zhang , Yulong Liu

A prescription is presented for a new and practical correlation coefficient, $\phi_K$, based on several refinements to Pearson's hypothesis test of independence of two variables. The combined features of $\phi_K$ form an advantage over…

Methodology · Statistics 2019-03-12 M. Baak , R. Koopman , H. Snoek , S. Klous

We introduce a new random matrix model called distance covariance matrix in this paper, whose normalized trace is equivalent to the distance covariance. We first derive a deterministic limit for the eigenvalue distribution of the distance…

Statistics Theory · Mathematics 2021-05-18 Weiming Li , Qinwen Wang , Jianfeng Yao

Independence screening is a variable selection method that uses a ranking criterion to select significant variables, particularly for statistical models with nonpolynomial dimensionality or "large p, small n" paradigms when p can be as…

Methodology · Statistics 2012-10-18 Gaorong Li , Heng Peng , Jun Zhang , Lixing Zhu

We consider the estimation of large covariance and precision matrices from high-dimensional sub-Gaussian or heavier-tailed observations with slowly decaying temporal dependence. The temporal dependence is allowed to be long-range so with…

Statistics Theory · Mathematics 2019-12-23 Hai Shu , Bin Nan

Building upon the Chatterjee correlation (2021: J. Am. Stat. Assoc. 116, p2009) for two real-valued variables, this study introduces a generalized measure of directed association between two vector variables, real or complex-valued, and of…

Methodology · Statistics 2024-06-26 Roberto D. Pascual-Marqui , Kieko Kochi , Toshihiko Kinoshita