中文
相关论文

相关论文: Post Selection Inference with Incomplete Maximum M…

200 篇论文

Recent advances in generative modeling have led to an increased interest in the study of statistical divergences as means of model comparison. Commonly used evaluation methods, such as the Frechet Inception Distance (FID), correlate well…

机器学习 · 统计学 2018-10-30 Mehdi S. M. Sajjadi , Olivier Bachem , Mario Lucic , Olivier Bousquet , Sylvain Gelly

Recent work has focused on the problem of nonparametric estimation of information divergence functionals. Many existing approaches are restrictive in their assumptions on the density support set or require difficult calculations at the…

信息论 · 计算机科学 2021-07-30 Kevin R. Moon , Kumar Sricharan , Kristjan Greenewald , Alfred O. Hero

Huge amount of applications in various fields, such as gene expression analysis or computer vision, undergo data sets with high-dimensional low-sample-size (HDLSS), which has putted forward great challenges for standard statistical and…

机器学习 · 计算机科学 2022-06-07 Liran Shen , Meng Joo Er , Qingbo Yin

In this article, we introduce a novel discrepancy called the maximum variance discrepancy for the purpose of measuring the difference between two distributions in Hilbert spaces that cannot be found via the maximum mean discrepancy. We also…

统计理论 · 数学 2020-12-08 Natsumi Makigusa

We introduce and study two new inferential challenges associated with the sequential detection of change in a high-dimensional mean vector. First, we seek a confidence interval for the changepoint, and second, we estimate the set of indices…

统计方法学 · 统计学 2023-03-03 Yudong Chen , Tengyao Wang , Richard J. Samworth

Mean embeddings provide an extremely flexible and powerful tool in machine learning and statistics to represent probability distributions and define a semi-metric (MMD, maximum mean discrepancy; also called N-distance or energy distance),…

机器学习 · 统计学 2019-05-17 Matthieu Lerasle , Zoltan Szabo , Timothee Mathieu , Guillaume Lecue

With the aim of generalizing histogram statistics to higher dimensional cases, density estimation via discrepancy based sequential partition (DSP) has been proposed to learn an adaptive piecewise constant approximation defined on a binary…

机器学习 · 统计学 2025-12-23 Zhengyang Lei , Lirong Qu , Sihong Shao , Yunfeng Xiong

The paper introduces a new kernel-based Maximum Mean Discrepancy (MMD) statistic for measuring the distance between two distributions given finitely-many multivariate samples. When the distributions are locally low-dimensional, the proposed…

机器学习 · 统计学 2018-09-03 Xiuyuan Cheng , Alexander Cloninger , Ronald R. Coifman

The problem of f-divergence estimation is important in the fields of machine learning, information theory, and statistics. While several nonparametric divergence estimators exist, relatively few have known convergence properties. In…

信息论 · 计算机科学 2015-03-16 Kevin R. Moon , Alfred O. Hero

Many machine learning applications such as in vision, biology and social networking deal with data in high dimensions. Feature selection is typically employed to select a subset of features which im- proves generalization accuracy as well…

机器学习 · 计算机科学 2016-06-15 Yamuna Prasad , Dinesh Khandelwal , K. K. Biswas

We propose a new sufficient dimension reduction approach designed deliberately for high-dimensional classification. This novel method is named maximal mean variance (MMV), inspired by the mean variance index first proposed by Cui, Li and…

统计方法学 · 统计学 2018-12-11 Xin Chen , Jingjing Wu , Zhigang Yao , Jia Zhang

Discrepancy measures between probability distributions are at the core of statistical inference and machine learning. In many applications, distributions of interest are supported on different spaces, and yet a meaningful correspondence…

机器学习 · 计算机科学 2021-11-23 Zhengxin Zhang , Youssef Mroueh , Ziv Goldfeld , Bharath K. Sriperumbudur

Semi-supervised (SS) inference has received much attention in recent years. Apart from a moderate-sized labeled data, L, the SS setting is characterized by an additional, much larger sized, unlabeled data, U. The setting of |U| >> |L|,…

统计方法学 · 统计学 2024-02-06 Yuqian Zhang , Abhishek Chakrabortty , Jelena Bradic

We consider the variable selection problem for two-sample tests, aiming to select the most informative variables to determine whether two collections of samples follow the same distribution. To address this, we propose a novel framework…

机器学习 · 统计学 2024-12-23 Jie Wang , Santanu S. Dey , Yao Xie

Multiple imputation (MI) has been widely applied to missing value problems in biomedical, social and econometric research, in order to avoid improper inference in the downstream data analysis. In the presence of high-dimensional data,…

统计方法学 · 统计学 2023-05-04 Zhiqi Bu , Zongyu Dai , Yiliang Zhang , Qi Long

A frequent problem in statistical science is how to properly handle missing data in matched paired observations. There is a large body of literature coping with the univariate case. Yet, the ongoing technological progress in measuring…

统计方法学 · 统计学 2022-06-06 Marcos Matabuena , Paulo Félix , Marc Ditzhaus , Juan Vidal , Francisco Gude

The selection of a validation basis from a full dataset is often required in industrial use of supervised machine learning algorithm. This validation basis will serve to realize an independent evaluation of the machine learning model. To…

机器学习 · 统计学 2021-04-30 Bertrand Iooss

We present a selective sampling method designed to accelerate the training of deep neural networks. To this end, we introduce a novel measurement, the minimal margin score (MMS), which measures the minimal amount of displacement an input…

机器学习 · 计算机科学 2019-11-19 Berry Weinstein , Shai Fine , Yacov Hel-Or

While sensitivity analysis improves the transparency and reliability of mathematical models, its uptake by modelers is still scarce. This is partially explained by its technical requirements, which may be hard to understand and implement by…

应用统计 · 统计学 2023-03-20 Arnald Puy , Pamphile T. Roy , Andrea Saltelli

We study the problem of detecting change points (CPs) that are characterized by a subset of dimensions in a multi-dimensional sequence. A method for detecting those CPs can be formulated as a two-stage method: one for selecting relevant…

机器学习 · 统计学 2018-03-05 Yuta Umezu , Ichiro Takeuchi