English
Related papers

Related papers: Model-free screening procedure for ultrahigh-dimen…

200 papers

We propose the Sobolev Independence Criterion (SIC), an interpretable dependency measure between a high dimensional random variable X and a response variable Y . SIC decomposes to the sum of feature importance scores and hence can be used…

Machine Learning · Computer Science 2019-11-01 Youssef Mroueh , Tom Sercu , Mattia Rigotti , Inkit Padhi , Cicero Dos Santos

High-dimensional data are commonly seen in modern statistical applications, variable selection methods play indispensable roles in identifying the critical features for scientific discoveries. Traditional best subset selection methods are…

Methodology · Statistics 2022-12-29 Tianzhou Ma , Hongjie Ke , Zhao Ren

Numerical modeling is essential for comprehending intricate physical phenomena in different domains. To handle complexity, sensitivity analysis, particularly screening, is crucial for identifying influential input parameters. Kernel-based…

Methodology · Statistics 2024-05-17 Guerlain Lambert , Céline Helbert , Claire Lauvernet

Many relations of scientific interest are nonlinear, and even in linear systems distributions are often non-Gaussian, for example in fMRI BOLD data. A class of search procedures for causal relations in high dimensional data relies on sample…

Artificial Intelligence · Computer Science 2014-01-30 Joseph D. Ramsey

The distribution-free method of conformal prediction (Vovk et al, 2005) has gained considerable attention in computer science, machine learning, and statistics. Candes et al. (2023) extended this method to right-censored survival data,…

Methodology · Statistics 2025-06-04 Jing Qin , Jin Piao , Jing Ning , Yu Shen

We provide a unified framework for independence and mean independence tests based on the Hilbert-Schmidt independence criterion, extending some previous results in the literature to hold in general topological spaces. We also present a…

Methodology · Statistics 2026-05-01 Daniel Diz-Castro , Manuel Febrero-Bande , Wenceslao González-Manteiga

The generalized linear models (GLM) have been widely used in practice to model non-Gaussian response variables. When the number of explanatory features is relatively large, scientific researchers are of interest to perform controlled…

Methodology · Statistics 2020-07-03 Chenguang Dai , Buyu Lin , Xin Xing , Jun S. Liu

A survival dataset describes a set of instances (e.g. patients) and provides, for each, either the time until an event (e.g. death), or the censoring time (e.g. when lost to follow-up - which is a lower bound on the time until the event).…

Machine Learning · Computer Science 2023-06-22 Ali Hossein Gharari Foomani , Michael Cooper , Russell Greiner , Rahul G. Krishnan

We propose a test of many zero parameter restrictions in a high dimensional linear iid regression model with $k$ $>>$ $n$ regressors. The test statistic is formed by estimating key parameters one at a time based on many low dimension…

Statistics Theory · Mathematics 2023-12-12 Jonathan B. Hill

Sliced inverse regression (SIR, Li 1991) is a pioneering work and the most recognized method in sufficient dimension reduction. While promising progress has been made in theory and methods of high-dimensional SIR, two remaining challenges…

Methodology · Statistics 2023-04-14 Qing Mai , Xiaofeng Shao , Runmin Wang , Xin Zhang

The statistical regression technique is an extraordinarily essential data fitting tool to explore the potential possible generation mechanism of the random phenomenon. Therefore, the model selection or the variable selection is becoming…

Methodology · Statistics 2020-03-25 Yue Su , Patrick Kandege Mwanakatwe

Chimeric antigen receptor T cell therapy has demonstrated innovative therapeutic effectiveness in fighting cancers; however, it is extremely expensive due to the intrinsic patient-to-patient variability in cell manufacturing. We propose in…

Quantitative Methods · Quantitative Biology 2021-05-18 Jialei Chen , Zhaonan Liu , Kan Wang , Chen Jiang , Chuck Zhang , Ben Wang

We emphasize that it is possible to improve the principle of unbiased risk estimation for model selection by addressing excess risk deviations in the design of penalization procedures. Indeed, we propose a modification of Akaike's…

Statistics Theory · Mathematics 2018-07-23 Adrien Saumard , Fabien Navarro

Causal inference has been increasingly reliant on observational studies with rich covariate information. To build tractable causal procedures, such as the doubly robust estimators, it is imperative to first extract important features from…

Methodology · Statistics 2022-02-08 Dingke Tang , Dehan Kong , Wenliang Pan , Linbo Wang

Gini distance correlation (GDC) was recently proposed to measure the dependence between a categorical variable, Y, and a numerical random vector, X. It mutually characterizes independence between X and Y. In this article, we utilize the GDC…

Methodology · Statistics 2023-04-19 Yongli Sang , Xin Dang

A stylized feature of high-dimensional data is that many variables have heavy tails, and robust statistical inference is critical for valid large-scale statistical inference. Yet, the existing developments such as Winsorization,…

Statistics Theory · Mathematics 2022-11-24 Jianqing Fan , Zhipeng Lou , Mengxin Yu

A model-agnostic variable importance method can be used with arbitrary prediction functions. Here we present some model-free methods that do not require access to the prediction function. This is useful when that function is proprietary and…

Machine Learning · Computer Science 2023-04-21 Naofumi Hama , Masayoshi Mase , Art B. Owen

Feature selection and reducing the dimensionality of data is an essential step in data analysis. In this work, we propose a new criterion for feature selection that is formulated as conditional information between features given the labeled…

Machine Learning · Statistics 2019-05-20 Salimeh Yasaei Sekeh , Alfred O. Hero

This paper presents a novel method for statistical inference in high-dimensional binary models with unspecified structure, where we leverage a (potentially misspecified) sparsity-constrained working generalized linear model (GLM) to…

Methodology · Statistics 2025-10-03 Xiaotian Hou , Peng Wang , Minge Xie , Linjun Zhang

Varying coefficient models have numerous applications in a wide scope of scientific areas. While enjoying nice interpretability, they also allow flexibility in modeling dynamic impacts of the covariates. But, in the new era of big data, it…

Methodology · Statistics 2014-10-27 Ming-Yen Cheng , Toshio Honda , Jin-Ting Zhang