English
Related papers

Related papers: A Kernel Method for the Two-Sample Problem

200 papers

We propose a novel method to detect and date structural breaks in the entire distribution of functional data. Theoretical guarantees are developed for our procedure under fewer assumptions than in the existing work. In particular, we…

Methodology · Statistics 2025-04-17 Peijun Sang , Bing Li

Regularized kernel methods such as support vector machines (SVM) and support vector regression (SVR) constitute a broad and flexible class of methods which are theoretically well investigated and commonly used in nonparametric…

Methodology · Statistics 2013-05-07 Robert Hable

This paper addresses the problem of regression to reconstruct functions, which are observed with superimposed errors at random locations. We address the problem in reproducing kernel Hilbert spaces. It is demonstrated that the estimator,…

Statistics Theory · Mathematics 2021-08-17 Paul Dommel , Alois Pichler

The reproducing kernel Hilbert space (RKHS) embedding of distributions offers a general and flexible framework for testing problems in arbitrary domains and has attracted considerable amount of attention in recent years. To gain insights…

Machine Learning · Statistics 2017-09-26 Krishnakumar Balasubramanian , Tong Li , Ming Yuan

We study two-sample variable selection: identifying variables that discriminate between the distributions of two sets of data vectors. Such variables help scientists understand the mechanisms behind dataset discrepancies. Although…

Machine Learning · Statistics 2025-11-06 Kensuke Mitsuzawa , Motonobu Kanagawa , Stefano Bortoli , Margherita Grossi , Paolo Papotti

We study the problems of sequential nonparametric two-sample and independence testing. Sequential tests process data online and allow using observed data to decide whether to stop and reject the null hypothesis or to collect more data,…

Machine Learning · Statistics 2023-07-21 Aleksandr Podkopaev , Aaditya Ramdas

We prove that kernel covariance embeddings lead to information-theoretically perfect separation of distinct continuous probability distributions. In statistical terms, we establish that testing for the \emph{equality} of two non-atomic…

Machine Learning · Statistics 2026-05-14 Leonardo V. Santoro , Kartik G. Waghmare , Victor M. Panaretos

This paper introduces algorithms to select/design kernels in Gaussian process regression/kriging surrogate modeling techniques. We adopt the setting of kernel method solutions in ad hoc functional spaces, namely Reproducing Kernel Hilbert…

Machine Learning · Statistics 2022-09-07 Jean-Luc Akian , Luc Bonnet , Houman Owhadi , Éric Savin

In this paper we suggest two statistical hypothesis tests for the regression function of binary classification based on conditional kernel mean embeddings. The regression function is a fundamental object in classification as it determines…

Machine Learning · Statistics 2022-06-22 Ambrus Tamás , Balázs Csanád Csáji

This paper generalizes regularized regression problems in a hyper-reproducing kernel Hilbert space (hyper-RKHS), illustrates its utility for kernel learning and out-of-sample extensions, and proves asymptotic convergence results for the…

Machine Learning · Computer Science 2022-10-20 Fanghui Liu , Lei Shi , Xiaolin Huang , Jie Yang , Johan A. K. Suykens

This paper addresses the problem of approximating an unknown function from point evaluations. When obtaining these point evaluations is costly, minimising the required sample size becomes crucial, and it is unreasonable to reserve a…

Numerical Analysis · Mathematics 2025-11-06 Nando Hegemann , Anthony Nouy , Philipp Trunschke

This paper provides a unifying view of optimal kernel hypothesis testing across the MMD two-sample, HSIC independence, and KSD goodness-of-fit frameworks. Minimax optimal separation rates in the kernel and $L^2$ metrics are presented, with…

Machine Learning · Statistics 2025-12-30 Antonin Schrab

Consider a setting with multiple units (e.g., individuals, cohorts, geographic locations) and outcomes (e.g., treatments, times, items), where the goal is to learn a multivariate distribution for each unit-outcome entry, such as the…

Machine Learning · Statistics 2025-10-21 Kyuseong Choi , Jacob Feitelberg , Caleb Chin , Anish Agarwal , Raaz Dwivedi

Motivated by the increasing use of kernel-based metrics for high-dimensional and large-scale data, we study the asymptotic behavior of kernel two-sample tests when the dimension and sample sizes both diverge to infinity. We focus on the…

Statistics Theory · Mathematics 2024-10-31 Jian Yan , Xianyang Zhang

Two-sample feature selection is the problem of finding features that describe a difference between two probability distributions, which is a ubiquitous problem in both scientific and engineering studies. However, existing methods have…

This paper studies convergence of empirical risks in reproducing kernel Hilbert spaces (RKHS). A conventional assumption in the existing research is that empirical training data do not contain any noise but this may not be satisfied in some…

Optimization and Control · Mathematics 2020-05-19 Shaoyan Guo , Huifu Xu , Liwei Zhang

Quantifying the difference between probability distributions is crucial in machine learning. However, estimating statistical divergences from empirical samples is challenging due to unknown underlying distributions. This work proposes the…

Machine Learning · Computer Science 2024-10-25 Jhoan K. Hoyos-Osorio , Luis G. Sanchez-Giraldo

We consider the random-design least-squares regression problem within the reproducing kernel Hilbert space (RKHS) framework. Given a stream of independent and identically distributed input/output data, we aim to learn a regression function…

Statistics Theory · Mathematics 2016-03-30 Aymeric Dieuleveut , Francis Bach

The kernel Maximum Mean Discrepancy~(MMD) is a popular multivariate distance metric between distributions that has found utility in two-sample testing. The usual kernel-MMD test statistic is a degenerate U-statistic under the null, and thus…

Methodology · Statistics 2025-09-16 Shubhanshu Shekhar , Ilmun Kim , Aaditya Ramdas

We review machine learning methods employing positive definite kernels. These methods formulate learning and estimation problems in a reproducing kernel Hilbert space (RKHS) of functions defined on the data domain, expanded in terms of a…

Statistics Theory · Mathematics 2009-09-29 Thomas Hofmann , Bernhard Schölkopf , Alexander J. Smola