中文
相关论文

相关论文: Distributed Conditional Feature Screening via Pear…

200 篇论文

The problem of selecting a handful of truly relevant variables in supervised machine learning algorithms is a challenging problem in terms of untestable assumptions that must hold and unavailability of theoretical assurances that selection…

统计方法学 · 统计学 2023-11-10 Mehdi Rostami , Olli Saarela

This paper proposes a model-free and data-adaptive feature screening method for ultra-high dimensional datasets. The proposed method is based on the projection correlation which measures the dependence between two random vectors. This…

统计方法学 · 统计学 2021-02-16 Wanjun Liu , Yuan Ke , Jingyuan Liu , Runze Li

Selecting relevant features associated with a given response variable is an important issue in many scientific fields. Quantifying quality and uncertainty of a selection result via false discovery rate (FDR) control has been of recent…

统计方法学 · 统计学 2020-12-17 Chenguang Dai , Buyu Lin , Xin Xing , Jun S. Liu

Feature screening is an important method to reduce the dimension and capture informative variables in ultrahigh-dimensional data analysis. Many methods have been developed for feature screening. These methods, however, are challenged by…

统计方法学 · 统计学 2019-01-08 Li-Pang Chen

Feature screening is a powerful tool in the analysis of high dimensional data. When the sample size $N$ and the number of features $p$ are both large, the implementation of classic screening methods can be numerically challenging. In this…

统计方法学 · 统计学 2019-03-12 Xingxiang Li , Runze Li , Zhiming Xia , Chen Xu

Variable selection has been widely used in data analysis for the past decades, and it becomes increasingly important in the Big Data era as there are usually hundreds of variables available in a dataset. To enhance interpretability of a…

统计方法学 · 统计学 2020-08-17 Yuxiang Xie , Kwun Chuen Gary Chan

Identifying multivariate dependencies in high-dimensional data is an important problem in large-scale inference. This problem has motivated recent advances in mining (partial) correlations, which focus on the challenging ultra-high…

统计方法学 · 统计学 2025-09-23 Emily Neo , Peter Radchenko , Bala Rajaratnam

Distributed optimization has been widely used as one of the most efficient approaches for model training with massive samples. However, large-scale learning problems with both massive samples and high-dimensional features widely exist in…

机器学习 · 计算机科学 2022-04-26 Runxue Bao , Xidong Wu , Wenhan Xian , Heng Huang

Controlling the false discovery rate (FDR) is a powerful approach to multiple testing. In many applications, the tested hypotheses have an inherent hierarchical structure. In this paper, we focus on the fixed sequence structure where the…

统计方法学 · 统计学 2016-11-11 Gavin Lynch , Wenge Guo , Sanat K. Sarkar , Helmut Finner

Controlling the False Discovery Rate (FDR) in a variable selection procedure is critical for reproducible discoveries, and it has been extensively studied in sparse linear models. However, it remains largely open in scenarios where the…

统计方法学 · 统计学 2023-11-16 Yang Cao , Xinwei Sun , Yuan Yao

In this era of "big" data, not only the large amount of data keeps motivating distributed computing, but concerns on data privacy also put forward the emphasis on distributed learning. To conduct feature selection and to control the false…

统计方法学 · 统计学 2020-08-11 Yu Gui

We develop a new class of distribution--free multiple testing rules for false discovery rate (FDR) control under general dependence. A key element in our proposal is a symmetrized data aggregation (SDA) approach to incorporating the…

统计方法学 · 统计学 2021-05-27 Lilun Du , Xu Guo , Wenguang Sun , Changliang Zou

Screening for ultrahigh dimensional features may encounter complicated issues such as outlying observations, heteroscedasticity or heavy-tailed distribution, multi-collinearity and confounding effects. Standard correlation-based marginal…

统计理论 · 数学 2018-12-27 Xiaochao Xia , Jialiang Li

Feature screening for ultrahigh-dimension, in general, proceeds with two essential steps. The first step is measuring and ranking the marginal dependence between response and covariates, and the second is determining the threshold. We…

统计方法学 · 统计学 2022-07-28 Linsui Deng , Yilin Zhang

This paper treats the problem of screening for variables with high correlations in high dimensional data in which there can be many fewer samples than variables. We focus on threshold-based correlation screening methods for three related…

机器学习 · 统计学 2015-03-18 Alfred O. Hero , Bala Rajaratnam

False discovery rate (FDR) control is a popular approach for maintaining the integrity of statistical analyses, especially in high-dimensional data settings, where multiple comparisons increase the risk of false positives. FDR control has…

信号处理 · 电气工程与系统科学 2026-03-03 Fabian Scheidt , Jasin Machkour , Michael Muma

False discovery rate (FDR) controlling procedures provide important statistical guarantees for the replicability in signal identification based on multiple hypotheses testing. In many fields of study, FDR controlling procedures are used in…

统计方法学 · 统计学 2022-10-04 Ran Dai , Cheng Zheng

We in this paper propose a directional regression based approach for ultrahigh dimensional sufficient variable screening with censored responses. The new method is designed in a model-free manner and thus can be adapted to various complex…

统计方法学 · 统计学 2018-02-28 Menghao Xu , Zhou Yu , Jun Shao

Multiple hypothesis testing is a fundamental problem in high dimensional inference, with wide applications in many scientific fields. In genome-wide association studies, tens of thousands of tests are performed simultaneously to find if any…

统计方法学 · 统计学 2010-12-21 Xu Han , Weijie Gu , Jianqing Fan

Controlling false discovery rate (FDR) is crucial for variable selection, multiple testing, among other signal detection problems. In literature, there is certainly no shortage of FDR control strategies when selecting individual features,…

统计方法学 · 统计学 2022-04-11 Jingyuan Liu , Ao Sun , Yuan Ke
‹ 上一页 1 2 3 10 下一页 ›