中文
相关论文

相关论文: Iterated Feature Screening based on Distance Corre…

200 篇论文

Feature screening for ultra high dimensional feature spaces plays a critical role in the analysis of data sets whose predictors exponentially exceed the number of observations. Such data sets are becoming increasingly prevalent in areas…

统计方法学 · 统计学 2018-01-31 Randall Reese , Xiaotian Dai , Guifang Fu

The paper considers variable selection in linear regression models where the number of covariates is possibly much larger than the number of observations. High dimensionality of the data brings in many complications, such as (possibly…

统计方法学 · 统计学 2016-11-29 Haeran Cho , Piotr Fryzlewicz

Modern bio-technologies have produced a vast amount of high-throughput data with the number of predictors far greater than the sample size. In order to identify more novel biomarkers and understand biological mechanisms, it is vital to…

机器学习 · 统计学 2018-05-18 Kevin He , Jian Kang , Hyokyoung Grace Hong , Ji Zhu , Yanming Li , Huazhen Lin , Han Xu , Yi Li

Feature selection is a critical step in the analysis of high-dimensional data, where the number of features often vastly exceeds the number of samples. Effective feature selection not only improves model performance and interpretability but…

机器学习 · 计算机科学 2025-01-27 Raquel Espinosa , Gracia Sánchez , José Palma , Fernando Jiménez

When the number of features exponentially outnumbers the number of samples, feature screening plays a pivotal role in reducing the dimension of the feature space and developing models based on such data. While most extant feature screening…

统计方法学 · 统计学 2019-11-19 Randall Reese

Variable selection is of increasing importance to address the difficulties of high dimensionality in many scientific areas. In this paper, we demonstrate a property for distance covariance, which is incorporated in a novel feature screening…

统计方法学 · 统计学 2014-09-03 Jing Kong , Sijian Wang , Grace Wahba

This paper treats the problem of screening for variables with high correlations in high dimensional data in which there can be many fewer samples than variables. We focus on threshold-based correlation screening methods for three related…

机器学习 · 统计学 2015-03-18 Alfred O. Hero , Bala Rajaratnam

In practical applications, one often does not know the "true" structure of the underlying conditional quantile function, especially in the ultra-high dimensional setting. To deal with ultra-high dimensionality, quantile-adaptive marginal…

统计方法学 · 统计学 2024-04-26 Daoji Li , Yinfei Kong , Dawit Zerom

In variable selection, most existing screening methods focus on marginal effects and ignore dependence between covariates. To improve the performance of selection, we incorporate pairwise effects in covariates for screening and…

统计方法学 · 统计学 2019-02-12 Siliang Gong , Kai Zhang , Yufeng Liu

Feature selection methods have an important role on the readability of data and the reduction of complexity of learning algorithms. In recent years, a variety of efforts are investigated on feature selection problems based on unsupervised…

机器学习 · 计算机科学 2019-12-12 Mohsen Ghassemi Parsa , Hadi Zare , Mehdi Ghatee

Many computer vision and medical imaging problems are faced with learning from large-scale datasets, with millions of observations and features. In this paper we propose a novel efficient learning scheme that tightens a sparsity constraint…

机器学习 · 统计学 2017-02-07 Adrian Barbu , Yiyuan She , Liangjing Ding , Gary Gramajo

Feature selection is frequently used as a pre-processing step to machine learning. It is a process of choosing a subset of original features so that the feature space is optimally reduced according to a certain evaluation criterion. The…

计算机视觉与模式识别 · 计算机科学 2014-01-07 Vijendra Singh , Shivani Pathak

Ultrahigh dimensional data sets are becoming increasingly prevalent in areas such as bioinformatics, medical imaging, and social network analysis. Sure independent screening of such data is commonly used to analyze such data. Nevertheless,…

统计方法学 · 统计学 2020-10-15 Randall Reese , Xiaotian Dai , Guifang Fu

Multi-view unsupervised feature selection has been proven to be efficient in reducing the dimensionality of multi-view unlabeled data with high dimensions. The previous methods assume all of the views are complete. However, in real…

机器学习 · 计算机科学 2023-01-02 Yanyong Huang , Kejun Guo , Xiuwen Yi , Zhong Li , Tianrui Li

Understanding how features interact with each other is of paramount importance in many scientific discoveries and contemporary applications. Yet interaction identification becomes challenging even for a moderate number of covariates. In…

统计方法学 · 统计学 2016-05-31 Yingying Fan , Yinfei Kong , Daoji Li , Jinchi Lv

High-dimensional sparse modeling with censored survival data is of great practical importance, and several methods have been proposed for variable selection based on different models. However, the impact of biased sample caused by…

统计方法学 · 统计学 2019-09-25 Li-Pang Chen

Gene expression datasets are usually of high dimensionality and therefore require efficient and effective methods for identifying the relative importance of their attributes. Due to the huge size of the search space of the possible…

机器学习 · 计算机科学 2022-06-10 Fernando Jiménez , Gracia Sánchez , José Palma , Luis Miralles-Pechuán , Juan Botía

Feature interactions can contribute to a large proportion of variation in many prediction models. In the era of big data, the coexistence of high dimensionality in both responses and covariates poses unprecedented challenges in identifying…

统计方法学 · 统计学 2016-05-12 Yinfei Kong , Daoji Li , Yingying Fan , Jinchi Lv

Ultra-high dimensional longitudinal data are increasingly common and the analysis is challenging both theoretically and methodologically. We offer a new automatic procedure for finding a sparse semivarying coefficient model, which is widely…

统计方法学 · 统计学 2014-09-24 Ming-Yen Cheng , Toshio Honda , Jialiang Li , Heng Peng

In recent years we have been able to gather large amounts of genomic data at a fast rate, creating situations where the number of variables greatly exceeds the number of observations. In these situations, most models that can handle a…

统计方法学 · 统计学 2025-02-07 Andrea Bratsberg , Abhik Ghosh , Magne Thoresen