English
Related papers

Related papers: Strong Sure Screening of Ultra-high Dimensional Da…

200 papers

This paper is concerned with the problems of interaction screening and nonlinear classification in a high-dimensional setting. We propose a two-step procedure, IIS-SQDA, where in the first step an innovated interaction screening (IIS)…

Machine Learning · Statistics 2015-06-04 Yingying Fan , Yinfei Kong , Daoji Li , Zemin Zheng

In modern drug development, the broader availability of high-dimensional observational data provides opportunities for scientist to explore subgroup heterogeneity, especially when randomized clinical trials are unavailable due to cost and…

Methodology · Statistics 2021-02-24 Xinzhou Guo , Linqing Wei , Chong Wu , Jingshen Wang

High-dimensional biomarkers such as genomics are increasingly being measured in randomized clinical trials. Consequently, there is a growing interest in developing methods that improve the power to detect biomarker-treatment interactions.…

Methodology · Statistics 2021-04-30 Jixiong Wang , Ashish Patel , James M. S. Wason , Paul J. Newcombe

Although much progress has been made in classification with high-dimensional features \citep{Fan_Fan:2008, JGuo:2010, CaiSun:2014, PRXu:2014}, classification with ultrahigh-dimensional features, wherein the features much outnumber the…

Machine Learning · Statistics 2016-11-14 Yanming Li , Hyokyoung Hong , Jian Kang , Kevin He , Ji Zhu , Yi Li

Ultrahigh-dimensional variable selection plays an increasingly important role in contemporary scientific discoveries and statistical research. Among others, Fan and Lv [J. R. Stat. Soc. Ser. B Stat. Methodol. 70 (2008) 849-911] propose an…

Methodology · Statistics 2012-11-14 Jianqing Fan , Rui Song

In practical applications, one often does not know the "true" structure of the underlying conditional quantile function, especially in the ultra-high dimensional setting. To deal with ultra-high dimensionality, quantile-adaptive marginal…

Methodology · Statistics 2024-04-26 Daoji Li , Yinfei Kong , Dawit Zerom

In recent years, deep learning has been at the center of analytics due to its impressive empirical success in analyzing complex data objects. Despite this success, most of the existing tools behave like black-box machines, thus the…

Machine Learning · Statistics 2022-11-02 Arkaprabha Ganguli , David Todem , Tapabrata Maiti

Advancement in technology has generated abundant high-dimensional data that allows integration of multiple relevant studies. Due to their huge computational advantage, variable screening methods based on marginal correlation have become…

Methodology · Statistics 2017-10-12 Tianzhou Ma , Zhao Ren , George C. Tseng

High-dimensional data are commonly seen in modern statistical applications, variable selection methods play indispensable roles in identifying the critical features for scientific discoveries. Traditional best subset selection methods are…

Methodology · Statistics 2022-12-29 Tianzhou Ma , Hongjie Ke , Zhao Ren

Independence screening is a powerful method for variable selection for `Big Data' when the number of variables is massive. Commonly used independence screening methods are based on marginal correlations or variations of it. In many…

Statistics Theory · Mathematics 2012-11-02 Emre Barut , Jianqing Fan , Anneleen Verhasselt

Gini distance correlation (GDC) was recently proposed to measure the dependence between a categorical variable, Y, and a numerical random vector, X. It mutually characterizes independence between X and Y. In this article, we utilize the GDC…

Methodology · Statistics 2023-04-19 Yongli Sang , Xin Dang

A daunting challenge faced by modern biological sciences is finding an efficient and computationally feasible approach to deal with the curse of high dimensionality. The problem becomes even more severe when the research focus is on…

Methodology · Statistics 2013-04-16 Hung Hung , Yu-Tin Lin , Pengwen Chen , Chen-Chien Wang , Su-Yun Huang , Jung-Ying Tzeng

Existing monitoring tools for multivariate data are often asymptotically distribution-free, computationally intensive, or require a large stretch of stable data. Many of these methods are not applicable to 'high dimension, low sample size'…

Methodology · Statistics 2023-05-12 Niladri Chakraborty , Chun Fai Lui , Ahmed Maged

Variable selection in high-dimensional space characterizes many contemporary problems in scientific discovery and decision making. Many frequently-used techniques are based on independence screening; examples include correlation ranking…

Methodology · Statistics 2008-12-18 Jianqing Fan , Richard Samworth , Yichao Wu

We propose a new model-free feature screening method based on energy distances for ultrahigh-dimensional binary classification problems. With a high probability, the proposed method retains only relevant features after discarding all the…

Methodology · Statistics 2023-05-19 Sarbojit Roy , Soham Sarkar , Subhajit Dutta , Anil K. Ghosh

The problem of interaction selection has recently caught much attention in high dimensional data analysis. This note aims to address and clarify several fundamental issues in interaction selection for linear regression models, especially…

Methodology · Statistics 2015-10-08 Ning Hao , Hao Helen Zhang

To date, testing interactions in high dimensions has been a challenging task. Existing methods often have issues with sensitivity to modeling assumptions and heavily asymptotic nominal p-values. To help alleviate these issues, we propose a…

Machine Learning · Statistics 2012-06-29 Noah Simon , Robert Tibshirani

Feature screening for ultrahigh-dimension, in general, proceeds with two essential steps. The first step is measuring and ranking the marginal dependence between response and covariates, and the second is determining the threshold. We…

Methodology · Statistics 2022-07-28 Linsui Deng , Yilin Zhang

A ubiquitous feature of data of our era is their extra-large sizes and dimensions. Analyzing such high-dimensional data poses significant challenges, since the feature dimension is often much larger than the sample size. This thesis…

Statistics Theory · Mathematics 2025-09-11 Kai Yang

We consider the problem of screening features in an ultrahigh-dimensional setting. Using maximum correlation, we develop a novel procedure called MC-SIS for feature screening, and show that MC-SIS possesses the sure screen property without…

Methodology · Statistics 2015-11-09 Qiming Huang , Yu Zhu