中文
相关论文

相关论文: Subsampling Winner Algorithm for Feature Selection…

200 篇论文

This paper provides a statistical analysis of high-dimensional batch Reinforcement Learning (RL) using sparse linear function approximation. When there is a large number of candidate features, our result sheds light on the fact that…

机器学习 · 计算机科学 2020-11-10 Botao Hao , Yaqi Duan , Tor Lattimore , Csaba Szepesvári , Mengdi Wang

In this work, we study and analyze different feature selection algorithms that can be used to classify cancer subtypes in case of highly varying high-dimensional data. We apply three different feature selection methods on five different…

机器学习 · 计算机科学 2021-10-01 Vaibhav Sinha , Siladitya Dash , Nazma Naskar , Sk Md Mosaddek Hossain

Principal component analysis (PCA) has been widely applied to dimensionality reduction and data pre-processing for different applications in engineering, biology and social science. Classical PCA and its variants seek for linear projections…

机器学习 · 计算机科学 2017-07-11 Xiaojun Chang , Feiping Nie , Yi Yang , Heng Huang

Feature selection is the process of identifying statistically most relevant features to improve the predictive capabilities of the classifiers. To find the best features subsets, the population based approaches like Particle Swarm…

神经与进化计算 · 计算机科学 2018-06-28 Naresh Mallenahalli , T. Hitendra Sarma

We consider the problem of variable selection in high-dimensional statistical models where the goal is to report a set of variables, out of many predictors $X_1, \dotsc, X_p$, that are relevant to a response of interest. For linear…

统计方法学 · 统计学 2019-03-20 Adel Javanmard , Hamid Javadi

Background: Identification of causal SNPs in most genome wide association studies relies on approaches that consider each SNP individually. However, there is a strong correlation structure among SNPs that need to be taken into account.…

应用统计 · 统计学 2012-11-02 Verena Zuber , A. Pedro Duarte Silva , Korbinian Strimmer

An important problem in bioinformatics is the inference of gene regulatory networks (GRN) from temporal expression profiles. In general, the main limitations faced by GRN inference methods is the small number of samples with huge…

计算机视觉与模式识别 · 计算机科学 2011-07-26 Fabrício Martins Lopes , David C. Martins-Jr , Junior Barrera , Roberto M. Cesar-Jr

This paper addresses the challenge of efficiently capturing a high proportion of true signals for subsequent data analyses when sample sizes are relatively limited with respect to data dimension. We propose the signal missing rate as a new…

统计方法学 · 统计学 2018-08-30 X. Jessie Jeng , Teng Zhang , Jung-Ying Tzeng

Significant pattern mining is a fundamental task in mining transactional data, requiring to identify patterns significantly associated with the value of a given feature, the target. In several applications, such as biomedicine, basket…

机器学习 · 计算机科学 2024-06-18 Leonardo Pellegrina , Fabio Vandin

Supervised matrix factorization (SMF) is a classical machine learning method that simultaneously seeks feature extraction and classification tasks, which are not necessarily a priori aligned objectives. Our goal is to use SMF to learn…

机器学习 · 统计学 2023-11-21 Joowon Lee , Hanbaek Lyu , Weixin Yao

Classification accuracy provided by a machine learning model depends a lot on the feature set used in the learning process. Feature Selection (FS) is an important and challenging pre-processing technique which helps to identify only the…

机器学习 · 计算机科学 2020-09-01 Ritam Guha , Manosij Ghosh , Shyok Mutsuddi , Ram Sarkar , Seyedali Mirjalili

High-dimensional clustering analysis is a challenging problem in statistics and machine learning, with broad applications such as the analysis of microarray data and RNA-seq data. In this paper, we propose a new clustering procedure called…

统计方法学 · 统计学 2022-10-31 Tianqi Liu , Yu Lu , Biqing Zhu , Hongyu Zhao

In this research, we have two serum SELDI (surface-enhanced laser desorption and ionization) mass spectra (MS) datasets to be used to select features amongst them to identify proteomic cancerous serums from normal serums. Features selection…

机器学习 · 计算机科学 2021-05-06 Ahmed Farag Seddik , Hassan Mostafa Ahmed

Stochastic First-Order (SFO) methods have been a cornerstone in addressing a broad spectrum of modern machine learning (ML) challenges. However, their efficacy is increasingly questioned, especially in large-scale applications where…

机器学习 · 计算机科学 2024-08-01 Di Zhang , Suvrajeet Sen

Feature screening is an important method to reduce the dimension and capture informative variables in ultrahigh-dimensional data analysis. Many methods have been developed for feature screening. These methods, however, are challenged by…

统计方法学 · 统计学 2019-01-08 Li-Pang Chen

One of the main obstacles of adopting digital pathology is the challenge of efficient processing of hyperdimensional digitized biopsy samples, called whole slide images (WSIs). Exploiting deep learning and introducing compact WSI…

计算机视觉与模式识别 · 计算机科学 2023-04-19 Azam Asilian Bidgoli , Shahryar Rahnamayan , Taher Dehkharghanian , Abtin Riasatian , H. R. Tizhoosh

Considering the case where the response variable is a categorical variable and the predictor is a random function, two novel functional sufficient dimensional reduction (FSDR) methods are proposed based on mutual information and square loss…

机器学习 · 统计学 2024-02-28 Xinyu Li , Jianjun Xu , Wenquan Cui , Haoyang Cheng

In today world of enormous amounts of data, it is very important to extract useful knowledge from it. This can be accomplished by feature subset selection. Feature subset selection is a method of selecting a minimum number of features with…

机器学习 · 计算机科学 2019-07-16 Agnip Dasgupta , Ardhendu Banerjee , Aniket Ghosh Dastidar , Antara Barman , Sanjay Chakraborty

Simultaneously performing variable selection and inference in high-dimensional regression models is an open challenge in statistics and machine learning. The increasing availability of vast amounts of variables requires the adoption of…

统计方法学 · 统计学 2025-05-08 Marco Molinari , Magne Thoresen

Gene expression data is widely used in disease analysis and cancer diagnosis. However, since gene expression data could contain thousands of genes simultaneously, successful microarray classification is rather difficult. Feature selection…

机器学习 · 计算机科学 2016-12-28 Li-Yeh Chuang , Chao-Hsuan Ke , Cheng-Hong Yang