中文
相关论文

相关论文: Learning Kernel Tests Without Data Splitting

200 篇论文

Stochastic processes are random variables with values in some space of paths. However, reducing a stochastic process to a path-valued random variable ignores its filtration, i.e. the flow of information carried by the process through time.…

Subsampling methods aim to select a subsample as a surrogate for the observed sample. Such methods have been used pervasively in large-scale data analytics, active learning, and privacy-preserving analysis in recent decades. Instead of…

机器学习 · 统计学 2022-06-03 Jingyi Zhang , Cheng Meng , Jun Yu , Mengrui Zhang , Wenxuan Zhong , Ping Ma

To adapt kernel two-sample and independence testing to complex structured data, aggregation of multiple kernels is frequently employed to boost testing power compared to single-kernel tests. However, we observe a phenomenon that directly…

机器学习 · 计算机科学 2025-10-14 Zhijian Zhou , Xunye Tian , Liuhua Peng , Chao Lei , Antonin Schrab , Danica J. Sutherland , Feng Liu

In this paper, we propose a test for the equality of multiple distributions based on kernel mean embeddings. Our framework provides a flexible way to handle multivariate or even high-dimensional data by virtue of kernel methods and allows…

统计理论 · 数学 2020-06-08 Ilmun Kim

Recent years have seen a surge in methods for two-sample testing, among which the Maximum Mean Discrepancy (MMD) test has emerged as an effective tool for handling complex and high-dimensional data. Despite its success and widespread…

机器学习 · 统计学 2026-05-21 Ikjun Choi , Ilmun Kim

Large scale online kernel learning aims to build an efficient and scalable kernel-based predictive model incrementally from a sequence of potentially infinite data points. A current key approach focuses on ways to produce an approximate…

机器学习 · 计算机科学 2019-09-25 Kai Ming Ting , Jonathan R. Wells , Takashi Washio

We consider the problem of two-sample testing in a semi-supervised setting with abundant unlabeled covariate data. Standard two-sample tests neglect covariate information, which has the potential to significantly boost performance. However,…

机器学习 · 统计学 2026-05-05 Gyumin Lee , Shubhanshu Shekhar , Ilmun Kim

Survival Analysis and Reliability Theory are concerned with the analysis of time-to-event data, in which observations correspond to waiting times until an event of interest such as death from a particular disease or failure of a component…

机器学习 · 统计学 2020-08-28 Tamara Fernandez , Nicolas Rivera , Wenkai Xu , Arthur Gretton

Spherical and hyperspherical data are commonly encountered in diverse applied research domains, underscoring the vital task of assessing independence within such data structures. In this context, we investigate the properties of test…

统计方法学 · 统计学 2024-01-23 Marija Cuparić , Bruno Ebner , Bojana Milošević

Kernel-based tests provide a simple yet effective framework that use the theory of reproducing kernel Hilbert spaces to design non-parametric testing procedures. In this paper we propose new theoretical tools that can be used to study the…

统计理论 · 数学 2022-09-02 Tamara Fernández , Nicolás Rivera

Machine learning and deep learning have been used extensively to classify physical surfaces through images and time-series contact data. However, these methods rely on human expertise and entail the time-consuming processes of data and…

机器学习 · 计算机科学 2023-08-10 Behnam Khojasteh , Friedrich Solowjow , Sebastian Trimpe , Katherine J. Kuchenbecker

Representations of probability measures in reproducing kernel Hilbert spaces provide a flexible framework for fully nonparametric hypothesis tests of independence, which can capture any type of departure from independence, including…

统计计算 · 统计学 2018-06-11 Qinyi Zhang , Sarah Filippi , Arthur Gretton , Dino Sejdinovic

Kernel-based hypothesis tests offer a flexible, non-parametric tool to detect high-order interactions in multivariate data, beyond pairwise relationships. Yet the scalability of such tests is limited by the computationally demanding…

统计方法学 · 统计学 2025-06-09 Zhaolu Liu , Robert L. Peach , Mauricio Barahona

Negative distance kernels $K(x,y) := - \|x-y\|$ were used in the definition of maximum mean discrepancies (MMDs) in statistics and lead to favorable numerical results in various applications. In particular, so-called slicing techniques for…

机器学习 · 统计学 2025-10-23 Nicolaj Rux , Michael Quellmalz , Gabriele Steidl

Two ubiquitous aspects of large-scale data analysis are that the data often have heavy-tailed properties and that diffusion-based or spectral-based methods are often used to identify and extract structure of interest. Perhaps surprisingly,…

机器学习 · 计算机科学 2010-05-11 Michael W. Mahoney , Hariharan Narayanan

Kernel quadrature is widely used to approximate integrals of smooth functions, with worst-case error typically decaying at the minimax rate $n^{-\alpha/d}$ for smoothness $\alpha$ in dimension $d$. Existing rate-optimal methods often depend…

统计计算 · 统计学 2026-05-19 Edoardo Bandoni , Christian Robert , Julien Stoehr

Change-point analysis plays a significant role in various fields to reveal discrepancies in distribution in a sequence of observations. While a number of algorithms have been proposed for high-dimensional data, kernel-based methods have not…

统计方法学 · 统计学 2023-01-10 Hoseung Song , Hao Chen

Knowing the error distribution is important in many multivariate time series applications. To alleviate the risk of error distribution mis-specification, testing methodologies are needed to detect whether the chosen error distribution is…

计量经济学 · 经济学 2020-08-04 Donghang Luo , Ke Zhu , Huan Gong , Dong Li

Distributional comparison is a fundamental problem in statistical data analysis with numerous applications in a variety of scientific and engineering fields. Numerous methods exist for distributional comparison but kernel Stein's method has…

统计理论 · 数学 2025-06-12 Xiaoda Qu , Baba C. Vemuri

Kernel mean embeddings are a popular tool that consists in representing probability measures by their infinite-dimensional mean embeddings in a reproducing kernel Hilbert space. When the kernel is characteristic, mean embeddings can be used…

机器学习 · 计算机科学 2021-06-29 Boris Muzellec , Francis Bach , Alessandro Rudi