English
Related papers

Related papers: Model-Free, Monotone Invariant and Computationally…

200 papers

Change point detection in high dimensional data has found considerable interest in recent years. Most of the literature either designs methodology for a retrospective analysis, where the whole sample is already available when the…

Statistics Theory · Mathematics 2020-12-16 Josua Gösmann , Christina Stoehr , Johannes Heiny , Holger Dette

This paper treats the problem of screening for variables with high correlations in high dimensional data in which there can be many fewer samples than variables. We focus on threshold-based correlation screening methods for three related…

Machine Learning · Statistics 2015-03-18 Alfred O. Hero , Bala Rajaratnam

Herein, we propose a Spearman rank correlation based screening procedure for ultrahigh-dimensional data with censored response case. The proposed method is model-free without specifying any regression forms of predictors or response…

Methodology · Statistics 2022-11-28 Hongni Wang , Jingxin Yan , Xiaodong Yan

We introduce a new class of methods for finite-sample false discovery rate (FDR) control in multiple testing problems with dependent test statistics where the dependence is fully or partially known. Our approach separately calibrates a…

Methodology · Statistics 2020-07-22 William Fithian , Lihua Lei

The problem of selecting a handful of truly relevant variables in supervised machine learning algorithms is a challenging problem in terms of untestable assumptions that must hold and unavailability of theoretical assurances that selection…

Methodology · Statistics 2023-11-10 Mehdi Rostami , Olli Saarela

Selecting relevant features associated with a given response variable is an important issue in many scientific fields. Quantifying quality and uncertainty of a selection result via false discovery rate (FDR) control has been of recent…

Methodology · Statistics 2020-12-17 Chenguang Dai , Buyu Lin , Xin Xing , Jun S. Liu

An important challenge in statistical analysis concerns the control of the finite sample bias of estimators. For example, the maximum likelihood estimator has a bias that can result in a significant inferential loss. This problem is…

Statistics Theory · Mathematics 2019-11-04 Stéphane Guerrier , Mucyo Karemera , Samuel Orso , Maria-Pia Victoria-Feser

Testing composite null hypotheses arises in various applications, such as mediation and replicability analyses. The problem becomes more challenging in high-throughput experiments where tens of thousands of features are examined…

Methodology · Statistics 2025-04-29 Pengfei Lyu , Xianyang Zhang , Hongyuan Cao

Large-scale multiple two-sample {\em Student}'s $t$ testing problems often arise from the statistical analysis of scientific data. To detect components with different values between two mean vectors, a well-known procedure is to apply the…

Methodology · Statistics 2014-10-17 Weidong Liu

We investigate the performance of a family of multiple comparison procedures for strong control of the False Discovery Rate ($\mathsf{FDR}$). The $\mathsf{FDR}$ is the expected False Discovery Proportion ($\mathsf{FDP}$), that is, the…

Statistics Theory · Mathematics 2008-11-21 Pierre Neuvial

We consider the problem of multiple hypothesis testing with generic side information: for each hypothesis $H_i$ we observe both a p-value $p_i$ and some predictor $x_i$ encoding contextual information about the hypothesis. For large-scale…

Methodology · Statistics 2018-07-26 Lihua Lei , William Fithian

Many important tasks of large-scale recommender systems can be naturally cast as testing multiple linear forms for noisy matrix completion. These problems, however, present unique challenges because of the subtle bias-and-variance tradeoff…

Methodology · Statistics 2025-03-12 Wanteng Ma , Lilun Du , Dong Xia , Ming Yuan

This paper is concerned with screening features in ultrahigh dimensional data analysis, which has become increasingly important in diverse scientific fields. We develop a sure independence screening procedure based on the distance…

Methodology · Statistics 2012-06-04 Runze Li , Wei Zhong , Liping Zhu

We consider the problem of screening features in an ultrahigh-dimensional setting. Using maximum correlation, we develop a novel procedure called MC-SIS for feature screening, and show that MC-SIS possesses the sure screen property without…

Methodology · Statistics 2015-11-09 Qiming Huang , Yu Zhu

This paper proposes a novel two-step strategy for testing the goodness-of-fit of parametric regression models in ultra-high dimensional sparse settings, where the predictor dimension far exceeds the sample size. This regime usually renders…

Methodology · Statistics 2025-12-30 Falong Tan , Jie Liu , Heng Peng , Lixing Zhu

We develop a new class of distribution--free multiple testing rules for false discovery rate (FDR) control under general dependence. A key element in our proposal is a symmetrized data aggregation (SDA) approach to incorporating the…

Methodology · Statistics 2021-05-27 Lilun Du , Xu Guo , Wenguang Sun , Changliang Zou

Variable screening has been a useful research area that deals with ultrahigh-dimensional data. When there exist both marginally and jointly dependent predictors to the response, existing methods such as conditional screening or iterative…

Methodology · Statistics 2023-07-10 Lei Fang , Qingcong Yuan , Xiangrong Yin , Chenglong Ye

A variable screening procedure via correlation learning was proposed Fan and Lv (2008) to reduce dimensionality in sparse ultra-high dimensional models. Even when the true model is linear, the marginal regression can be highly nonlinear. To…

Methodology · Statistics 2011-01-19 Jianqing Fan , Yang Feng , Rui Song

Controlling the False Discovery Rate (FDR) in a variable selection procedure is critical for reproducible discoveries, and it has been extensively studied in sparse linear models. However, it remains largely open in scenarios where the…

Methodology · Statistics 2023-11-16 Yang Cao , Xinwei Sun , Yuan Yao

Feature screening for ultra high dimensional feature spaces plays a critical role in the analysis of data sets whose predictors exponentially exceed the number of observations. Such data sets are becoming increasingly prevalent in areas…

Methodology · Statistics 2018-01-31 Randall Reese , Xiaotian Dai , Guifang Fu