中文
相关论文

相关论文: Distributed Conditional Feature Screening via Pear…

200 篇论文

The applications of traditional statistical feature selection methods to high-dimension, low sample-size data often struggle and encounter challenging problems, such as overfitting, curse of dimensionality, computational infeasibility, and…

机器学习 · 统计学 2023-12-19 Kexuan Li , Fangfang Wang , Lingli Yang , Ruiqi Liu

While data-driven confounder selection requires careful consideration, it is frequently employed in observational studies. Widely recognized criteria for confounder selection include the minimal-set approach, which involves selecting…

统计方法学 · 统计学 2025-08-21 Kazuharu Harada , Masataka Taguri

In many applications, the process of identifying a specific feature of interest often involves testing multiple hypotheses for their joint statistical significance. Examples include mediation analysis which simultaneously examines the…

统计方法学 · 统计学 2023-05-30 Linsui Deng , Kejun He , Xianyang Zhang

High-dimensional data are commonly seen in modern statistical applications, variable selection methods play indispensable roles in identifying the critical features for scientific discoveries. Traditional best subset selection methods are…

统计方法学 · 统计学 2022-12-29 Tianzhou Ma , Hongjie Ke , Zhao Ren

Variable screening is a fast dimension reduction technique for assisting high dimensional feature selection. As a preselection method, it selects a moderate size subset of candidate variables for further refining via feature selection to…

统计理论 · 数学 2015-06-09 Xiangyu Wang , Chenlei Leng , David B. Dunson

This paper is concerned with screening features in ultrahigh dimensional data analysis, which has become increasingly important in diverse scientific fields. We develop a sure independence screening procedure based on the distance…

统计方法学 · 统计学 2012-06-04 Runze Li , Wei Zhong , Liping Zhu

In large-scale biomedical research, it's common to gather ultra-high dimensional data that includes right-censored survival times. Feature screening has emerged as a crucial statistical technique for handling such data. In this paper, we…

统计方法学 · 统计学 2026-03-31 Shuya Chen , Heng Peng , Min Zhou

Controlling the false discovery rate (FDR) is a popular approach to multiple testing, variable selection, and related problems of simultaneous inference. In many contemporary applications, models are not specified by discrete variables,…

统计理论 · 数学 2024-04-16 Mateo Díaz , Venkat Chandrasekaran

Feature screening is useful and popular to detect informative predictors for ultrahigh-dimensional data before developing proceeding statistical analysis or constructing statistical models. While a large body of feature screening procedures…

统计方法学 · 统计学 2020-08-12 Li-Pang Chen

Feature selection methods are widely used to address the high computational overheads and curse of dimensionality in classifying high-dimensional data. Most conventional feature selection methods focus on handling homogeneous features,…

机器学习 · 计算机科学 2021-11-17 Xuyang Yan , Mrinmoy Sarkar , Biniam Gebru , Shabnam Nazmi , Abdollah Homaifar

This paper develops a framework for testing for associations in a possibly high-dimensional linear model where the number of features/variables may far exceed the number of observational units. In this framework, the observations are split…

统计方法学 · 统计学 2018-05-04 Rina Foygel Barber , Emmanuel J. Candes

Feature screening approaches are effective in selecting active features from data with ultrahigh dimensionality and increasing complexity; however, the majority of existing feature screening approaches are either restricted to a univariate…

统计方法学 · 统计学 2023-05-09 Shaofei Zhao , Guifang Fu

Testing for differences in features between clusters in various applications often leads to inflated false positives when practitioners use the same dataset to identify clusters and then test features, an issue commonly known as ``double…

统计方法学 · 统计学 2024-10-10 Lijun Wang , Yingxin Lin , Hongyu Zhao

In this study, we propose a method Distributionally Robust Safe Screening (DRSS), for identifying unnecessary samples and features within a DR covariate shift setting. This method effectively combines DR learning, a paradigm aimed at…

Herein, we propose a Spearman rank correlation based screening procedure for ultrahigh-dimensional data with censored response case. The proposed method is model-free without specifying any regression forms of predictors or response…

统计方法学 · 统计学 2022-11-28 Hongni Wang , Jingxin Yan , Xiaodong Yan

This paper proposes a new feature screening method for the multi-response ultrahigh dimensional linear model by empirical likelihood. Through a multivariate moment condition, the empirical likelihood induced ranking statistics can exploit…

统计方法学 · 统计学 2022-06-07 Jun Lu , Qinqin Hu , Lu Lin

Multiple hypothesis testing is a fundamental problem in high dimensional inference, with wide applications in many scientific fields. In genome-wide association studies, tens of thousands of tests are performed simultaneously to find if any…

统计方法学 · 统计学 2011-11-16 Jianqing Fan , Xu Han , Weijie Gu

In many fields of science, we observe a response variable together with a large number of potential explanatory variables, and would like to be able to discover which variables are truly associated with the response. At the same time, we…

统计方法学 · 统计学 2015-10-15 Rina Foygel Barber , Emmanuel J. Candès

In big data analysis, a simple task such as linear regression can become very challenging as the variable dimension $p$ grows. As a result, variable screening is inevitable in many scientific studies. In recent years, randomized algorithms…

统计方法学 · 统计学 2019-02-13 Yu-Hsiang Cheng , Tzee-Ming Huang , Su-Yun Huang

Independence screening is a variable selection method that uses a ranking criterion to select significant variables, particularly for statistical models with nonpolynomial dimensionality or "large p, small n" paradigms when p can be as…

统计方法学 · 统计学 2012-10-18 Gaorong Li , Heng Peng , Jun Zhang , Lixing Zhu