English
Related papers

Related papers: Stratified incomplete local simplex tests for curv…

200 papers

In this paper, we introduce the \textbf{G}eneralized \textbf{L}inear \textbf{S}pectral \textbf{S}tatistics (GLSS) of a high-dimensional sample covariance matrix $\bm{S}_n$, denoted as $\operatorname{tr}f(\bm{S}_n)\bm{B}_n$, which…

Statistics Theory · Mathematics 2025-12-03 Yanlin Hu , Qing Yang , Xiao Han

Model selection and sparse recovery are two important problems for which many regularization methods have been proposed. We study the properties of regularization methods in both problems under the unified framework of regularized least…

Statistics Theory · Mathematics 2009-09-03 Jinchi Lv , Yingying Fan

Within the nonparametric regression model with unknown regression function $l$ and independent, symmetric errors, a new multiscale signed rank statistic is introduced and a conditional multiple test of the simple hypothesis $l=0$ against a…

Statistics Theory · Mathematics 2008-12-18 Angelika Rohde

Partial Least Squares (PLS) methods have been heavily exploited to analyse the association between two blocs of data. These powerful approaches can be applied to data sets where the number of variables is greater than the number of…

Machine Learning · Statistics 2017-02-24 Pierre Lafaye de Micheaux , Benoit Liquet , Matthew Sutton

In classification problems, sampling bias between training data and testing data is critical to the ranking performance of classification scores. Such bias can be both unintentionally introduced by data collection and intentionally…

Methodology · Statistics 2017-11-02 Chandler Zuo

We consider the multiple testing of the general regression framework aiming at studying the relationship between a univariate response and a p-dimensional predictor. To test the hypothesis of the effect of each predictor, we construct an…

Methodology · Statistics 2023-06-19 Zhigen Zhao , Xin Xing

Stratified sampling is a fast and simple method to generate point sets with uniform distribution in hypercubes. However, for the most common paraxial stratfication it has the prominent drawback that the number of sampled points in n…

Computation · Statistics 2018-06-14 Simon Wessing

Recent works have proposed optimal subsampling algorithms to improve computational efficiency in large datasets and to design validation studies in the presence of measurement error. Existing approaches generally fall into two categories:…

Methodology · Statistics 2025-12-25 Jasper B. Yang , Thomas Lumley , Bryan E. Shepherd , Pamela A. Shaw

We propose a nonparametric bootstrap procedure for two-phase stratified sampling without replacement. In this design, a weighted likelihood estimator is known to have smaller asymptotic variance than under the convenient assumption of…

Statistics Theory · Mathematics 2014-09-26 Takumi Saegusa

In this paper, we propose novel, fully Bayesian non-parametric tests for one-sample and two-sample multivariate location problems. We model the underlying distribution using a Dirichlet process prior, and develop a testing procedure based…

Statistics Theory · Mathematics 2021-08-03 Indrabati Bhattacharya , Subhashis Ghosal

Imbalanced classification and spurious correlation are common challenges in data science and machine learning. Both issues are linked to data imbalance, with certain groups of data samples significantly underrepresented, which in turn would…

Machine Learning · Statistics 2026-02-10 Ryumei Nakada , Yichen Xu , Lexin Li , Linjun Zhang

This article studies global testing of the slope function in functional linear regression model in the framework of reproducing kernel Hilbert space. We propose a new testing statistic based on smoothness regularization estimators. The…

Statistics Theory · Mathematics 2021-10-13 Jianjun Xu , Wenquan Cui

A variety of statistics based on sample spacings has been studied in the literature for testing goodness-of-fit to parametric distributions. To test the goodness-of-fit to a nonparametric class of univariate shape-constrained densities,…

Statistics Theory · Mathematics 2024-10-28 Kwun Chuen Gary Chan , Hok Kan Ling , Chuan-Fa Tang , Sheung Chi Phillip Yam

For data with high-dimensional covariates but small to moderate sample sizes, the analysis of single datasets often generates unsatisfactory results. The integrative analysis of multiple independent datasets provides an effective way of…

Methodology · Statistics 2015-01-19 Yuan Huang , Qingzhao Zhang , Sanguo Zhang , Jian Huang , Shuangge Ma

We develop a unified $L$-statistic testing framework for high-dimensional regression coefficients that adapts to unknown sparsity. The proposed statistics rank coordinate-wise evidence measures and aggregate the top $k$ signals, bridging…

Applications · Statistics 2026-02-10 Ping Zhao , Fengyi Song , Huifang Ma

The classical iteratively reweighted least-squares (IRLS) algorithm aims to recover an unknown signal from linear measurements by performing a sequence of weighted least squares problems, where the weights are recursively updated at each…

Machine Learning · Statistics 2024-06-06 Chiraag Kaushik , Justin Romberg , Vidya Muthukumar

We consider the problem of variable screening in ultra-high dimensional generalized linear models (GLMs) of non-polynomial orders. Since the popular SIS approach is extremely unstable in the presence of contamination and noise, we discuss a…

Statistics Theory · Mathematics 2022-11-15 Abhik Ghosh , Erica Ponzi , Torkjel Sandanger , Magne Thoresen

Theoretical works on supervised transfer learning (STL) -- where the learner has access to labeled samples from both source and target distributions -- have for the most part focused on statistical aspects of the problem, while efficient…

Machine Learning · Statistics 2025-07-08 Yuyang Deng , Samory Kpotufe

We consider the problem of learning a high-dimensional multi-task regression model, under sparsity constraints induced by presence of grouping structures on the input covariates and on the output predictors. This problem is primarily…

Machine Learning · Statistics 2012-05-10 Seunghak Lee , Eric P. Xing

A significant hurdle for analyzing large sample data is the lack of effective statistical computing and inference methods. An emerging powerful approach for analyzing large sample data is subsampling, by which one takes a random subsample…

Methodology · Statistics 2015-11-24 Rong Zhu , Ping Ma , Michael W. Mahoney , Bin Yu