English
Related papers

Related papers: A sure independence screening procedure for ultra-…

200 papers

This paper studies the distributed conditional feature screening for massive data with ultrahigh-dimensional features. Specifically, three distributed partial correlation feature screening methods (SAPS, ACPS and JDPS methods) are firstly…

Methodology · Statistics 2024-03-12 Naiwen Pang , Xiaochao Xia

This paper treats the problem of screening for variables with high correlations in high dimensional data in which there can be many fewer samples than variables. We focus on threshold-based correlation screening methods for three related…

Machine Learning · Statistics 2015-03-18 Alfred O. Hero , Bala Rajaratnam

Many applications benefit from theory relevant to the identification of variables having large correlations or partial correlations in high dimension. Recently there has been progress in the ultra-high dimensional setting when the sample…

Statistics Theory · Mathematics 2022-08-25 Yun Wei , Bala Rajaratnam , Alfred O. Hero

Machine learning methods are used to discover complex nonlinear relationships in biological and medical data. However, sophisticated learning models are computationally unfeasible for data with millions of features. Here we introduce the…

In many applied fields incomplete covariate vectors are commonly encountered. It is well known that this can be problematic when making inference on model parameters, but its impact on prediction performance is less understood. We develop a…

Methodology · Statistics 2020-07-14 Garritt L. Page , Fernando A. Quintana , Peter Müller

We consider generalized linear regression analysis with left-censored covariate due to the lower limit of detection. Complete case analysis by eliminating observations with values below limit of detection yields valid estimates for…

Methodology · Statistics 2014-12-09 Shengchun Kong , Bin Nan

Understanding variable dependence, particularly eliciting their statistical properties given a set of covariates, provides the mathematical foundation in practical operations management such as risk analysis and decision-making given…

Methodology · Statistics 2023-09-06 Yunyun Wang , Tatsushi Oka , Dan Zhu

Statistical inference in high dimensional settings has recently attracted enormous attention within the literature. However, most published work focuses on the parametric linear regression problem. This paper considers an important…

Methodology · Statistics 2019-11-14 Qi Gao , Randy C. S. Lai , Thomas C. M. Lee , Yao Li

In this paper, we develop invariance-based procedures for testing and inference in high-dimensional regression models. These procedures, also known as randomization tests, provide several important advantages. First, for the global null…

Methodology · Statistics 2023-12-27 Wenxuan Guo , Panos Toulis

The paper considers variable selection in linear regression models where the number of covariates is possibly much larger than the number of observations. High dimensionality of the data brings in many complications, such as (possibly…

Methodology · Statistics 2016-11-29 Haeran Cho , Piotr Fryzlewicz

This paper proposes a new method for estimating high-dimensional binary choice models. We consider a semiparametric model that places no distributional assumptions on the error term, allows for heteroskedastic errors, and permits endogenous…

Econometrics · Economics 2025-07-15 Fu Ouyang , Thomas Tao Yang

Missing data are often dealt with multiple imputation. A crucial part of the multiple imputation process is selecting sensible models to generate plausible values for incomplete data. A method based on posterior predictive checking is…

Computation · Statistics 2026-05-14 Mingyang Cai , Stef van Buuren , Gerko Vink

Identifying important biomarkers that are predictive for cancer patients' prognosis is key in gaining better insights into the biological influences on the disease and has become a critical component of precision medicine. The emergence of…

Methodology · Statistics 2016-03-22 Hyokyoung Grace Hong , Jian Kang , Yi Li

We propose a two-step procedure to detect cointegration in high-dimensional settings, focusing on sparse relationships. First, we use the adaptive LASSO to identify the small subset of integrated covariates driving the equilibrium…

Methodology · Statistics 2026-03-05 Jesus Gonzalo , Jean-Yves Pitarakis

We introduce a simple diagnostic test for assessing the overall or partial goodness of fit of a linear causal model with errors being independent of the covariates. In particular, we consider situations where hidden confounding is…

Methodology · Statistics 2023-03-06 Christoph Schultheiss , Peter Bühlmann , Ming Yuan

Modern high-throughput biomedical devices routinely produce data on a large scale, and the analysis of high-dimensional datasets has become commonplace in biomedical studies. However, given thousands or tens of thousands of measured…

Methodology · Statistics 2022-02-28 Vladimir Vutov , Thorsten Dickhaus

Test of independence is of fundamental importance in modern data analysis, with broad applications in variable selection, graphical models, and causal inference. When the data is high dimensional and the potential dependence signal is…

Methodology · Statistics 2023-06-13 Zhanrui Cai , Jing Lei , Kathryn Roeder

This paper investigates the utilization of maximum and average distance correlations for multivariate independence testing. We characterize their consistency properties in high-dimensional settings with respect to the number of marginally…

Machine Learning · Statistics 2025-06-11 Cencheng Shen , Yuexiao Dong

An endeavor central to precision medicine is predictive biomarker discovery; they define patient subpopulations which stand to benefit most, or least, from a given treatment. The identification of these biomarkers is often the byproduct of…

Methodology · Statistics 2022-08-01 Philippe Boileau , Nina Ting Qi , Mark J. van der Laan , Sandrine Dudoit , Ning Leng

Semiparametric regression offers a flexible framework for modeling non-linear relationships between a response and covariates. A prime example are generalized additive models where splines (say) are used to approximate non-linear functional…

Statistics Theory · Mathematics 2018-10-05 Francis K. C. Hui , Chong You , Han Lin Shang , Samuel Müller