中文
相关论文

相关论文: Iterated Feature Screening based on Distance Corre…

200 篇论文

Feature selection from a large number of covariates (aka features) in a regression analysis remains a challenge in data science, especially in terms of its potential of scaling to ever-enlarging data and finding a group of scientifically…

机器学习 · 统计学 2020-02-10 Yiying Fan , Jiayang Sun

Individualized treatment rules can lead to better health outcomes when patients have heterogeneous responses to treatment. Very few individualized treatment rule estimation methods are compatible with a multi-treatment observational study…

统计方法学 · 统计学 2019-11-14 Owen E. Leete , Nathan Kallus , Michael G. Hudgens , Sonia Napravnik , Michael R. Kosorok

During the last decades, learning a low-dimensional space with discriminative information for dimension reduction (DR) has gained a surge of interest. However, it's not accessible for these DR methods to achieve satisfactory performance…

机器学习 · 计算机科学 2019-11-19 Xiangzhu Meng , Huibing Wang , Lin Feng

Feature selection is an important but challenging task in causal inference for obtaining unbiased estimates of causal quantities. Properly selected features in causal inference not only significantly reduce the time required to implement a…

统计方法学 · 统计学 2025-02-04 Tianyu Yang , Md. Noor-E-Alam

Along with the flourish of the information age, massive amounts of data are generated day by day. Due to the large-scale and high-dimensional characteristics of these data, it is often difficult to achieve better decision-making in…

机器学习 · 计算机科学 2023-04-04 Peican Zhu , Xin Hou , Keke Tang , Zhen Wang , Feiping Nie

This thesis responds to the challenges of using a large number, such as thousands, of features in regression and classification problems. There are two situations where such high dimensional features arise. One is when high dimensional…

机器学习 · 统计学 2007-09-20 Longhai Li

This paper is concerned with screening features in ultrahigh dimensional data analysis, which has become increasingly important in diverse scientific fields. We develop a sure independence screening procedure based on the distance…

统计方法学 · 统计学 2012-06-04 Runze Li , Wei Zhong , Liping Zhu

We propose a flexible nonparametric regression method for ultrahigh-dimensional data. As a first step, we propose a fast screening method based on the favored smoothing bandwidth of the marginal local constant regression. Then, an iterative…

统计方法学 · 统计学 2018-07-30 Yang Feng , Yichao Wu , Leonard Stefanski

In the field of data mining, how to deal with high-dimensional data is an inevitable problem. Unsupervised feature selection has attracted more and more attention because it does not rely on labels. The performance of spectral-based…

机器学习 · 计算机科学 2021-01-01 Zhengxin Li , Feiping Nie , Jintang Bian , Xuelong Li

Estimating a sparse covariance matrix is a fundamental problem in high-dimensional statistics. However, thresholding methods developed for independent data are generally not directly applicable to high-dimensional time series, where…

统计方法学 · 统计学 2026-05-15 Wenhao Zhang , Zhaoxing Gao

An analysis of high-dimensional data can offer a detailed description of a system but is often challenged by the curse of dimensionality. General dimensionality reduction techniques can alleviate such difficulty by extracting a few…

统计方法学 · 统计学 2021-09-28 Di Bo , Hoon Hwangbo , Vinit Sharma , Corey Arndt , Stephanie C. TerMaath

The accelerated failure time (AFT) models have proved useful in many contexts, though heavy censoring (as for example in cancer survival) and high dimensionality (as for example in microarray data) cause difficulties for model fitting and…

统计方法学 · 统计学 2013-12-10 Md Hasinur Rahaman Khan , J. Ewart H. Shaw

Gini distance correlation (GDC) was recently proposed to measure the dependence between a categorical variable, Y, and a numerical random vector, X. It mutually characterizes independence between X and Y. In this article, we utilize the GDC…

统计方法学 · 统计学 2023-04-19 Yongli Sang , Xin Dang

Feature selection (FS) is a process which attempts to select more informative features. In some cases, too many redundant or irrelevant features may overpower main features for classification. Feature selection can remedy this problem and…

机器学习 · 计算机科学 2013-06-07 A. Nisthana Parveen , H. Hannah Inbarani , E. N. Sathishkumar

We introduce a new approach to variable selection, called Predictive Correlation Screening, for predictor design. Predictive Correlation Screening (PCS) implements false positive control on the selected variables, is well suited to small…

机器学习 · 统计学 2013-04-11 Hamed Firouzi , Bala Rajaratnam , Alfred Hero

Recent studies try to use hyperspectral imaging (HSI) to detect foreign matters in products because it enables to visualize the invisible wavelengths including ultraviolet and infrared. Considering the enormous image channels of the HSI,…

计算机视觉与模式识别 · 计算机科学 2024-01-10 Dongeon Kim , YeongHyeon Park

In hyperspectral remote sensing data mining, it is important to take into account of both spectral and spatial information, such as the spectral signature, texture feature and morphological property, to improve the performances, e.g., the…

计算机视觉与模式识别 · 计算机科学 2019-04-09 Lefei Zhang , Qian Zhang , Bo Du , Xin Huang , Yuan Yan Tang , Dacheng Tao

In this paper, we study a novel approach for the estimation of quantiles when facing potential right censoring of the responses. Contrary to the existing literature on the subject, the adopted strategy of this paper is to tackle censoring…

统计方法学 · 统计学 2017-03-24 Mickaël De Backer , Anouar El Ghouch , Ingrid Van Keilegom

Variable selection plays a fundamental role in high-dimensional data analysis. Various methods have been developed for variable selection in recent years. Well-known examples are forward stepwise regression (FSR) and least angle regression…

统计方法学 · 统计学 2018-02-01 Siliang Gong , Kai Zhang , Yufeng Liu

We present a conformal inference method for constructing lower prediction bounds for survival times from right-censored data, extending recent approaches designed for more restrictive type-I censoring scenarios. The proposed method imputes…

统计方法学 · 统计学 2025-05-26 Matteo Sesia , Vladimir Svetnik