English
Related papers

Related papers: A note on marginal correlation based screening

200 papers

Detecting dependence between variables is a crucial issue in statistical science. In this paper, we propose a novel metric called label projection correlation to measure the dependence between numerical and categorical variables. The…

Methodology · Statistics 2025-06-24 Yixiao Liu , Pengjian Shang

Simple correlation coefficients between two variables have been generalized to measure association between two matrices in many ways. Coefficients such as the RV coefficient, the distance covariance (dCov) coefficient and kernel based…

Methodology · Statistics 2014-08-19 Julie Josse , Susan Holmes

Decision trees and their ensembles are endowed with a rich set of diagnostic tools for ranking and screening variables in a predictive model. Despite the widespread use of tree based variable importance measures, pinning down their…

Machine Learning · Statistics 2020-12-14 Jason M. Klusowski , Peter M. Tian

Confounding matters in almost all observational studies that focus on causality. In order to eliminate bias caused by connfounders, oftentimes a substantial number of features need to be collected in the analysis. In this case, large p…

Statistics Theory · Mathematics 2019-12-30 Shinyuu Lee , Yuru Zhu

Given data obtained under two sampling conditions, it is often of interest to identify variables that behave differently in one condition than in the other. We introduce a method for differential analysis of second-order behavior called…

Methodology · Statistics 2016-02-26 Kelly Bodwin , Kai Zhang , Andrew Nobel

We consider the problem of screening features in an ultrahigh-dimensional setting. Using maximum correlation, we develop a novel procedure called MC-SIS for feature screening, and show that MC-SIS possesses the sure screen property without…

Methodology · Statistics 2015-11-09 Qiming Huang , Yu Zhu

We study the data-driven selection of causal graphical models using constraint-based algorithms, which determine the existence or non-existence of edges (causal connections) in a graph based on testing a series of conditional independence…

Methodology · Statistics 2026-04-29 Daniel Malinsky

High-dimensional biomarkers such as genomics are increasingly being measured in randomized clinical trials. Consequently, there is a growing interest in developing methods that improve the power to detect biomarker-treatment interactions.…

Methodology · Statistics 2021-04-30 Jixiong Wang , Ashish Patel , James M. S. Wason , Paul J. Newcombe

Spurious correlations allow flexible models to predict well during training but poorly on related test distributions. Recent work has shown that models that satisfy particular independencies involving correlation-inducing \textit{nuisance}…

Machine Learning · Computer Science 2022-06-10 Mark Goldstein , Jörn-Henrik Jacobsen , Olina Chau , Adriel Saporta , Aahlad Puli , Rajesh Ranganath , Andrew C. Miller

Two popular variable screening methods under the ultra-high dimensional setting with the desirable sure screening property are the sure independence screening (SIS) and the forward regression (FR). Both are classical variable screening…

Methodology · Statistics 2015-11-05 Ming-Yen Cheng , Sanying Feng , Gaorong Li , Heng Lian

Mendelian randomization uses genetic variants to make causal inferences about the effect of a risk factor on an outcome. With fine-mapped genetic data, there may be hundreds of genetic variants in a single gene region any of which could be…

Methodology · Statistics 2017-07-10 Stephen Burgess , Verena Zuber , Elsa Valdes-Marquez , Benjamin B Sun , Jemma C Hopewell

Feature screening is useful and popular to detect informative predictors for ultrahigh-dimensional data before developing proceeding statistical analysis or constructing statistical models. While a large body of feature screening procedures…

Methodology · Statistics 2020-08-12 Li-Pang Chen

Identifying how dependence relationships vary across different conditions plays a significant role in many scientific investigations. For example, it is important for the comparison of biological systems to see if relationships between…

Methodology · Statistics 2023-07-31 Hoseung Song , Michael C. Wu

Although conceptually related, variable selection and relative importance (RI) analysis have been treated quite differently in the literature. While RI is typically used for post-hoc model explanation, this paper explores its potential for…

Machine Learning · Statistics 2026-04-24 Tien-En Chang , Argon Chen

In this paper we propose a novel variable selection method for two-view settings, or for vector-valued supervised learning problems. Our framework is able to handle extremely large scale selection tasks, where number of data samples could…

Machine Learning · Computer Science 2023-07-06 Sandor Szedmak , Riikka Huusari , Tat Hong Duong Le , Juho Rousu

Over the last couple of decades, several copula based methods have been proposed in the literature to test for the independence among several random variables. But these existing tests are not invariant under monotone transformations of the…

Statistics Theory · Mathematics 2019-11-15 Angshuman Roy , Anil Ghosh , Alok Goswami , C. A. Murthy

A simple and intuitive method for feature selection consists of choosing the feature subset that maximizes a nonparametric measure of dependence between the response and the features. A popular proposal from the literature uses the…

Machine Learning · Statistics 2024-06-12 Keli Liu , Feng Ruan

Despite the availability of numerous statistical and machine learning tools for joint feature modeling, many scientists investigate features marginally, i.e., one feature at a time. This is partly due to training and convention but also…

Methodology · Statistics 2021-12-01 Jingyi Jessica Li , Yiling Chen , Xin Tong

High-dimensional data are commonly seen in modern statistical applications, variable selection methods play indispensable roles in identifying the critical features for scientific discoveries. Traditional best subset selection methods are…

Methodology · Statistics 2022-12-29 Tianzhou Ma , Hongjie Ke , Zhao Ren

We propose a test of independence of two multivariate random vectors, given a sample from the underlying population. Our approach, which we call MINT, is based on the estimation of mutual information, whose decomposition into joint and…

Methodology · Statistics 2017-11-20 Thomas B. Berrett , Richard J. Samworth
‹ Prev 1 3 4 5 6 7 10 Next ›