English
Related papers

Related papers: A fast algorithm for computing distance correlatio…

200 papers

Performance estimation under covariate shift is a crucial component of safe AI model deployment, especially for sensitive use-cases. Recently, several solutions were proposed to tackle this problem, most leveraging model predictions or…

Computer Vision and Pattern Recognition · Computer Science 2025-04-11 Mélanie Roschewitz , Ben Glocker

In this paper numerical methods of computing distances between two Radon measures on R are discussed. Efficient algorithms for Wasserstein-type metrics are provided. In particular, we propose a novel algorithm to compute the flat metric…

Numerical Analysis · Mathematics 2013-04-15 Jedrzej Jablonski , Anna Marciniak-Czochra

The discrepancy between two independent samples \(X_1,\dots,X_n\) and \(Y_1,\dots,Y_n\) drawn from the same distribution on $\mathbb{R}^d$ typically has order \(O(\sqrt{n})\) even in one dimension. We give a simple online algorithm that…

Probability · Mathematics 2026-01-23 Gleb Smirnov , Roman Vershynin

We provide a simple method and relevant theoretical analysis for efficiently estimating higher-order lp distances. While the analysis mainly focuses on l4, our methodology extends naturally to p = 6,8,10..., (i.e., when p is even).…

Machine Learning · Computer Science 2012-03-19 Ping Li , Michael W. Mahoney , Yiyuan She

We propose a novel method for testing serial independence of object-valued time series in metric spaces, which is more general than Euclidean or Hilbert spaces. The proposed method is fully nonparametric, free of tuning parameters, and can…

Methodology · Statistics 2023-07-31 Feiyu Jiang , Hanjia Gao , Xiaofeng Shao

We study the problem of high-dimensional sparse mean estimation in the presence of an $\epsilon$-fraction of adversarial outliers. Prior work obtained sample and computationally efficient algorithms for this task for identity-covariance…

Data Structures and Algorithms · Computer Science 2024-07-08 Ilias Diakonikolas , Daniel M. Kane , Sushrut Karmalkar , Ankit Pensia , Thanasis Pittas

For a bivariate time series $((X_i,Y_i))_{i=1,...,n}$ we want to detect whether the correlation between $X_i$ and $Y_i$ stays constant for all $i = 1,...,n$. We propose a nonparametric change-point test statistic based on Kendall's tau and…

Statistics Theory · Mathematics 2022-04-12 Herold Dehling , Daniel Vogel , Martin Wendler , Dominik Wied

We give the first polynomial time algorithm for \emph{list-decodable covariance estimation}. For any $\alpha > 0$, our algorithm takes input a sample $Y \subseteq \mathbb{R}^d$ of size $n\geq d^{\mathsf{poly}(1/\alpha)}$ obtained by…

Data Structures and Algorithms · Computer Science 2022-06-23 Misha Ivkov , Pravesh K. Kothari

Properly estimating correlations between objects at different spatial scales necessitates $\mathcal{O}(n^2)$ distance calculations. For this reason, most widely adopted packages for estimating correlations use clustering algorithms to…

Re-sampling based statistical tests are known to be computationally heavy, but reliable when small sample sizes are available. Despite their nice theoretical properties not much effort has been put to make them efficient. In this paper we…

Methodology · Statistics 2018-06-29 Christina Chatzipantsiou , Marios Dimitriadis , Manos Papadakis , Michail Tsagris

We propose a methodology to explore and measure the pairwise correlations that exist between variables in a dataset. The methodology leverages copulas for encoding dependence between two variables, state-of-the-art optimal transport for…

Machine Learning · Statistics 2016-11-01 Gautier Marti , Sebastien Andler , Frank Nielsen , Philippe Donnat

Sampling from multiple distributions so as to maximize overlap has been studied by statisticians since the 1950s. Since the 2000s, such correlated sampling from the probability simplex has been a powerful building block in disparate areas…

Data Structures and Algorithms · Computer Science 2025-11-18 Joseph , Naor , Nitya Raju , Abhishek Shetty , Aravind Srinivasan , Renata Valieva , David Wajc

We consider the detection problem of correlations in a $p$-dimensional Gaussian vector, when we observe $n$ independent, identically distributed random vectors, for $n$ and $p$ large. We assume that the covariance matrix varies in some…

Statistics Theory · Mathematics 2016-01-27 Cristina Butucea , Rania Zgheib

Pairwise likelihood is a useful approximation to the full likelihood function for covariance estimation in high-dimensional context. It simplifies high-dimensional dependencies by combining marginal bivariate likelihood objects, thus making…

Methodology · Statistics 2024-07-25 Alessandro Casa , Davide Ferrari , Zhendong Huang

Identifying dependency between two random variables is a fundamental problem. The clear interpretability and ability of a procedure to provide information on the form of possible dependence is particularly important when exploring…

Methodology · Statistics 2026-04-27 Bogdan Ćmiel , Teresa Ledwina

Reliably measuring the collinearity of bivariate data is crucial in statistics, particularly for time-series analysis or ongoing studies in which incoming observations can significantly impact current collinearity estimates. Leveraging…

Methodology · Statistics 2024-06-11 Marc Harary

Correlation and spectral analysis represent the standard tools to study interdependence in statistical data. However, for the stochastic processes with heavy-tailed distributions such that the variance diverges, these tools are inadequate.…

Statistical Mechanics · Physics 2015-06-22 Agnieszka Wyłomańska , Aleksei Chechkin , Janusz Gajda , Igor M. Sokolov

Distance correlation has gained much recent attention in the data science community: the sample statistic is straightforward to compute and asymptotically equals zero if and only if independence, making it an ideal choice to discover any…

Machine Learning · Statistics 2024-06-27 Cencheng Shen , Sambit Panda , Joshua T. Vogelstein

We present new algorithms for estimating and testing \emph{collision probability}, a fundamental measure of the spread of a discrete distribution that is widely used in many scientific fields. We describe an algorithm that satisfies…

Machine Learning · Statistics 2025-04-21 Robert Busa-Fekete , Umar Syed

Robust covariance estimation is the following, well-studied problem in high dimensional statistics: given $N$ samples from a $d$-dimensional Gaussian $\mathcal{N}(\boldsymbol{0}, \Sigma)$, but where an $\varepsilon$-fraction of the samples…

Data Structures and Algorithms · Computer Science 2020-06-25 Jerry Li , Guanghao Ye