English
Related papers

Related papers: Kernel Two-Sample and Independence Tests for Non-S…

200 papers

Conditional independence (CI) is central to causal inference, feature selection, and graphical modeling, yet it is untestable in many settings without additional assumptions. Existing CI tests often rely on restrictive structural…

Machine Learning · Computer Science 2025-12-23 Alek Frohlich , Vladimir Kostic , Karim Lounici , Daniel Perazzo , Massimiliano Pontil

This article addresses the problem of testing the conditional independence of two generic random vectors $X$ and $Y$ given a third random vector $Z$, which plays an important role in statistical and machine learning applications. We propose…

Methodology · Statistics 2024-07-26 Yi Zhang , Linjun Huang , Yun Yang , Xiaofeng Shao

Hilbert-Schmidt independence criterion and distance covariance are methods to describe independence of random variables using either the Kronecker product of positive definite kernels or the Kronecker product of conditionally negative…

Functional Analysis · Mathematics 2022-01-05 Jean Carlo Guella

We describe a novel non-parametric statistical hypothesis test of relative dependence between a source variable and two candidate target variables. Such a test enables us to determine whether one source variable is significantly more…

Machine Learning · Statistics 2015-05-28 Wacha Bounliphone , Arthur Gretton , Arthur Tenenhaus , Matthew Blaschko

Improvement of statistical learning models in order to increase efficiency in solving classification or regression problems is still a goal pursued by the scientific community. In this way, the support vector machine model is one of the…

Machine Learning · Statistics 2019-11-22 Anderson Ara , Mateus Maia , Samuel Macêdo , Francisco Louzada

Measures of discrepancy between probability distributions (statistical distance) are widely used in the fields of artificial intelligence and machine learning. We describe how certain measures of statistical distance can be implemented as…

Accelerator Physics · Physics 2022-12-21 Chad E. Mitchell , Robert D. Ryne , Kilean Hwang

We introduce kernel nonparametric tests for Lancaster three-variable interaction and for total independence, using embeddings of signed measures into a reproducing kernel Hilbert space. The resulting test statistics are straightforward to…

Methodology · Statistics 2013-06-11 Dino Sejdinovic , Arthur Gretton , Wicher Bergsma

Identifying relationships among stochastic processes is a core objective in many fields, such as economics. While the standard toolkit for multivariate time series analysis has many advantages, it can be difficult to capture nonlinear…

Methodology · Statistics 2026-05-06 Michael Wieck-Sosa , Michel F. C. Haddad , Aaditya Ramdas

In order to fully utilize "big data", it is often required to use "big models". Such models tend to grow with the complexity and size of the training data, and do not make strong parametric assumptions upfront on the nature of the…

Machine Learning · Statistics 2015-04-17 Vikas Sindhwani , Haim Avron

Conditional mean independence (CMI) testing is crucial for statistical tasks including model determination and variable importance evaluation. In this work, we introduce a novel population CMI measure and a bootstrap-based testing procedure…

Machine Learning · Statistics 2025-01-30 Yi Zhang , Linjun Huang , Yun Yang , Xiaofeng Shao

This paper proposes a new mutual independence test for a large number of high dimensional random vectors. The test statistic is based on the characteristic function of the empirical spectral distribution of the sample covariance matrix. The…

Statistics Theory · Mathematics 2012-05-31 G. M. Pan , J. Gao , Y. Yang , M. Guo

A new goodness-of-fit test for normality in high-dimension (and Reproducing Kernel Hilbert Space) is proposed. It shares common ideas with the Maximum Mean Discrepancy (MMD) it outperforms both in terms of computation time and applicability…

Statistics Theory · Mathematics 2014-04-14 Jérémie Kellner , Alain Celisse

The distribution closeness testing (DCT) assesses whether the distance between a distribution pair is at least $\epsilon$-far. Existing DCT methods mainly measure discrepancies between a distribution pair defined on discrete one-dimensional…

Machine Learning · Computer Science 2025-10-10 Zhijian Zhou , Liuhua Peng , Xunye Tian , Feng Liu

The application of Gaussian processes (GPs) to large data sets is limited due to heavy memory and computational requirements. A variety of methods has been proposed to enable scalability, one of which is to exploit structure in the kernel…

Machine Learning · Computer Science 2019-12-30 Jan Graßhoff , Alexandra Jankowski , Philipp Rostalski

In the statistical literature, as well as in artificial intelligence and machine learning, measures of discrepancy between two probability distributions are largely used to develop measures of goodness-of-fit. We concentrate on quadratic…

Methodology · Statistics 2025-10-01 Marianthi Markatou , Giovanni Saraceno

We propose robust two-sample tests for comparing means in time series. The framework accommodates a wide range of applications, including structural breaks, treatment-control comparisons, and group-averaged panel data. We first consider…

Econometrics · Economics 2025-12-23 Ulrich Hounyo , Min Seong Kim

Building spatial process models that capture nonstationary behavior while delivering computationally efficient inference is challenging. Nonstationary spatially varying kernels (see, e.g., Paciorek, 2003) offer flexibility and richness, but…

Methodology · Statistics 2025-07-01 Sébastien Coube-Sisqueille , Sudipto Banerjee , Benoît Liquet

In this paper, we consider the problem of testing independence in high-dimensional settings with missing data. Building upon a recently proposed Kendall-based statistic, we introduce two new modifications specifically designed to…

Methodology · Statistics 2026-04-28 Marija Cuparić , Bojana Milošević , Jelena Radojević

We consider the problem of causal structure learning in the setting of heterogeneous populations, i.e., populations in which a single causal structure does not adequately represent all population members, as is common in biological and…

Machine Learning · Statistics 2022-02-21 Alex Markham , Richeek Das , Moritz Grosse-Wentrup

Domain specific (dis-)similarity or proximity measures used e.g. in alignment algorithms of sequence data, are popular to analyze complex data objects and to cover domain specific data properties. Without an underlying vector space these…

Data Structures and Algorithms · Computer Science 2014-11-07 Andrej Gisbrecht , Frank-Michael Schleif