中文
相关论文

相关论文: A Normality Test for High-dimensional Data based o…

200 篇论文

We consider the problem of subspace estimation in situations where the number of available snapshots and the observation dimension are comparable in magnitude. In this context, traditional subspace methods tend to fail because the…

信息论 · 计算机科学 2016-11-15 Pascal Vallet , Philippe Loubaton , Xavier Mestre

Testing the equality of the covariance matrices of two high-dimensional samples is a fundamental inference problem in statistics. Several tests have been proposed but they are either too liberal or too conservative when the required…

统计理论 · 数学 2023-01-04 Jin-Ting Zhang , Jingyi Wang , Tianming Zhu

We propose a novel non-parametric adaptive anomaly detection algorithm for high dimensional data based on score functions derived from nearest neighbor graphs on $n$-point nominal data. Anomalies are declared whenever the score of a test…

机器学习 · 计算机科学 2009-10-29 Manqi Zhao , Venkatesh Saligrama

We suggest a robust nearest-neighbor approach to classifying high-dimensional data. The method enhances sensitivity by employing a threshold and truncates to a sequence of zeros and ones in order to reduce the deleterious impact of…

统计理论 · 数学 2009-09-02 Yao-ban Chan , Peter Hall

Estimating some mathematical expectations from partially observed data and in particular missing outcomes is a central problem encountered in numerous fields such as transfer learning, counterfactual analysis or causal inference. Matching…

统计理论 · 数学 2025-05-01 Simon Viel , Lionel Truquet , Ikko Yamane

Two-sample tests are important areas aiming to determine whether two collections of observations follow the same distribution or not. We propose two-sample tests based on integral probability metric (IPM) for high-dimensional samples…

机器学习 · 统计学 2023-04-21 Jie Wang , Minshuo Chen , Tuo Zhao , Wenjing Liao , Yao Xie

Maximum Mean Discrepancy (MMD) has been widely used in the areas of machine learning and statistics to quantify the distance between two distributions in the $p$-dimensional Euclidean space. The asymptotic property of the sample MMD has…

统计理论 · 数学 2023-08-29 Hanjia Gao , Xiaofeng Shao

Multivariate pattern analyses approaches in neuroimaging are fundamentally concerned with investigating the quantity and type of information processed by various regions of the human brain; typically, estimates of classification accuracy…

机器学习 · 统计学 2016-10-11 Charles Y. Zheng , Yuval Benjamini

This paper presents a new similarity measure to be used for general tasks including supervised learning, which is represented by the K-nearest neighbor classifier (KNN). The proposed similarity measure is invariant to large differences in…

机器学习 · 计算机科学 2014-09-04 Ahmad Basheer Hassanat

In this paper, for the problem of heteroskedastic general linear hypothesis testing (GLHT) in high-dimensional settings, we propose a random integration method based on the reference L2-norm to deal with such problems. The asymptotic…

统计理论 · 数学 2024-09-19 Mingxiang Cao , Hongwei Zhang , Kai Xu , Daojiang He

This paper investigates the utilization of maximum and average distance correlations for multivariate independence testing. We characterize their consistency properties in high-dimensional settings with respect to the number of marginally…

机器学习 · 统计学 2025-06-11 Cencheng Shen , Yuexiao Dong

We consider testing for two-sample means of high dimensional populations by thresholding. Two tests are investigated, which are designed for better power performance when the two population mean vectors differ only in sparsely populated…

统计方法学 · 统计学 2014-10-13 Song Xi Chen , Jun Li , Ping-Shou Zhong

Hypothesis testing in high dimensional data is a notoriously difficult problem without direct access to competing models' likelihood functions. This paper argues that statistical divergences can be used to quantify the difference between…

数据分析、统计与概率 · 物理学 2024-08-02 Jeremy J. H. Wilkinson , Christopher G. Lester

Many modern methods for prediction leverage nearest neighbor search to find past training examples most similar to a test example, an idea that dates back in text to at least the 11th century and has stood the test of time. This monograph…

机器学习 · 计算机科学 2025-02-25 George H. Chen , Devavrat Shah

To assess whether there is some signal in a big database, aggregate tests for the global null hypothesis of no effect are routinely applied in practice before more specialized analysis is carried out. Although a plethora of aggregate tests…

统计理论 · 数学 2024-05-08 Anders Bredahl Kock , David Preinerstorfer

Nonparametric generalized likelihood ratio test is popularly used for model checking for regressions. However, there are two issues that may be the barriers for its powerfulness. First, the bias term in its liming null distribution causes…

统计方法学 · 统计学 2015-07-23 Cuizhen Niu , Xu Guo , Lixing Zhu

Let $\mathbf{X} = (X_i)_{1\leq i \leq n}$ be an i.i.d. sample of square-integrable variables in $\mathbb{R}^d$, \GB{with common expectation $\mu$ and covariance matrix $\Sigma$, both unknown.} We consider the problem of testing if $\mu$ is…

机器学习 · 计算机科学 2021-10-11 Gilles Blanchard , Jean-Baptiste Fermanian

This paper introduces the generalized Hausman test as a novel method for detecting non-normality of the latent variable distribution of unidimensional Item Response Theory (IRT) models for binary data. The test utilizes the pairwise maximum…

统计方法学 · 统计学 2024-02-14 Lucia Guastadisegni , Silvia Cagnone , Irini Moustaki , Vassilis Vasdekis

We propose novel methodology for testing equality of model parameters between two high-dimensional populations. The technique is very general and applicable to a wide range of models. The method is based on sample splitting: the data is…

统计方法学 · 统计学 2013-01-17 Nicolas Städler , Sach Mukherjee

When data is of an extraordinarily large size or physically stored in different locations, the distributed nearest neighbor (NN) classifier is an attractive tool for classification. We propose a novel distributed adaptive NN classifier for…

机器学习 · 统计学 2023-06-06 Ruiqi Liu , Ganggang Xu , Zuofeng Shang