中文
相关论文

相关论文: A Normality Test for High-dimensional Data based o…

200 篇论文

We develop an asymptotic theory for $L^2$ norms of sample mean vectors of high-dimensional data. An invariance principle for the $L^2$ norms is derived under conditions that involve a delicate interplay between the dimension $p$, the sample…

统计理论 · 数学 2015-03-13 Mengyu Xu , Danna Zhang , Wei Biao Wu

The test of independence is a crucial component of modern data analysis. However, traditional methods often struggle with the complex dependency structures found in high-dimensional data. To overcome this challenge, we introduce a novel…

统计方法学 · 统计学 2024-09-13 Mingshuo Liu , Doudou Zhou , Hao Chen

Anomaly and similarity detection in multidimensional series have a long history and have found practical usage in many different fields such as medicine, networks, and finance. Anomaly detection is of great appeal for many different…

统计计算 · 统计学 2012-05-10 Paolo D'Alberto , Chris Drome , Ali Dasdan

Missing values are a common phenomenon in all areas of applied research. While various imputation methods are available for metrically scaled variables, methods for categorical data are scarce. An imputation method that has been shown to…

统计方法学 · 统计学 2017-10-04 Shahla Faisal , Gerhard Tutz

Motivated by the likelihood ratio test under the Gaussian assumption, we develop a maximum sum-of-squares test for conducting hypothesis testing on high dimensional mean vector. The proposed test which incorporates the dependence among the…

统计方法学 · 统计学 2015-10-21 Xianyang Zhang

Fitting high-dimensional statistical models often requires the use of non-linear parameter estimation procedures. As a consequence, it is generally impossible to obtain an exact characterization of the probability distribution of the…

统计方法学 · 统计学 2014-04-03 Adel Javanmard , Andrea Montanari

Nearest neighbor imputation is popular for handling item nonresponse in survey sampling. In this article, we study the asymptotic properties of the nearest neighbor imputation estimator for general population parameters, including…

统计方法学 · 统计学 2017-07-05 Shu Yang , Jae Kwang Kim

Principal component analysis continues to be a powerful tool in dimension reduction of high dimensional data. We assume a variance-diverging model and use the high-dimension, low-sample-size asymptotics to show that even though the…

统计理论 · 数学 2020-09-28 Sungkyu Jung

Estimating the intrinsic dimensionality (ID) of data is a fundamental problem in machine learning and computer vision, providing insight into the true degrees of freedom underlying high-dimensional observations. Existing methods often rely…

机器学习 · 计算机科学 2026-03-12 Eng-Jon Ong , Omer Bobrowski , Gesine Reinert , Primoz Skraba

We consider the estimation of densities in multiple subpopulations, where the available sample size in each subpopulation greatly varies. This problem occurs in epidemiology, for example, where different diseases may share similar…

统计方法学 · 统计学 2021-09-15 Jiaming Qiu , Xiongtao Dai , Zhengyuan Zhu

This article deals with the analysis of high dimensional data that come from multiple sources (experiments) and thus have different possibly correlated responses, but share the same set of predictors. The measurements of the predictors may…

统计方法学 · 统计学 2020-07-01 Guorong Dai , Ursula U. Müller , Raymond J. Carroll

Uncertainty estimation aims to evaluate the confidence of a trained deep neural network. However, existing uncertainty estimation approaches rely on low-dimensional distributional assumptions and thus suffer from the high dimensionality of…

机器学习 · 计算机科学 2023-10-26 Tsai Hor Chan , Kin Wai Lau , Jiajun Shen , Guosheng Yin , Lequan Yu

In the high dimensional regression analysis when the number of predictors is much larger than the sample size, an important question is to select the important variable which are relevant to the response variable of interest. Variable…

统计方法学 · 统计学 2023-01-09 Pengsheng Ji , Zhigen Zhao

Thanks to its favorable properties, the multivariate normal distribution is still largely employed for modeling phenomena in various scientific fields. However, when the number of components $p$ is of the same asymptotic order as the sample…

统计理论 · 数学 2022-11-17 Caizhu Huang , Claudia Di Caterina , Nicola Sartori

A classifier for two or more samples is proposed when the data are high-dimensional and the underlying distributions may be non-normal. The classifier is constructed as a linear combination of two easily computable and interpretable…

统计理论 · 数学 2016-08-02 M. Rauf Ahmad , Tatjana Pavlenko

In this paper, we develop invariance-based procedures for testing and inference in high-dimensional regression models. These procedures, also known as randomization tests, provide several important advantages. First, for the global null…

统计方法学 · 统计学 2023-12-27 Wenxuan Guo , Panos Toulis

Manifold hypothesis states that data points in high-dimensional space actually lie in close vicinity of a manifold of much lower dimension. In many cases this hypothesis was empirically verified and used to enhance unsupervised and…

We propose a novel and computationally efficient approach for nonparametric conditional density estimation in high-dimensional settings that achieves dimension reduction without imposing restrictive distributional or functional form…

计量经济学 · 经济学 2025-10-14 Jianhua Mei , Fu Ouyang , Thomas T. Yang

Cognitive diagnosis models have been popularly used in fields such as education, psychology, and social sciences. While parametric likelihood estimation is a prevailing method for fitting cognitive diagnosis models, nonparametric…

统计理论 · 数学 2025-10-01 Chengyu Cui , Yanlong Liu , Gongjun Xu

Nearest neighbor methods are a popular class of nonparametric estimators with several desirable properties, such as adaptivity to different distance scales in different regions of space. Prior work on convergence rates for nearest neighbor…

机器学习 · 计算机科学 2014-07-03 Kamalika Chaudhuri , Sanjoy Dasgupta