English
Related papers

Related papers: Model-free screening procedure for ultrahigh-dimen…

200 papers

Feature or variable selection is a problem inherent to large data sets. While many methods have been proposed to deal with this problem, some can scale poorly with the number of predictors in a data set. Screening methods scale linearly…

Methodology · Statistics 2023-01-09 Naveed Merchant , Jeffrey D. Hart

Conventional survival metrics, such as Harrell's concordance index (CI) and the Brier Score, rely on the independent censoring assumption for valid inference with right-censored data. However, in the presence of so-called dependent…

Machine Learning · Statistics 2025-05-20 Christian Marius Lillelund , Shi-ang Qi , Russell Greiner

In nonparametric independence testing, we observe i.i.d.\ data $\{(X_i,Y_i)\}_{i=1}^n$, where $X \in \mathcal{X}, Y \in \mathcal{Y}$ lie in any general spaces, and we wish to test the null that $X$ is independent of $Y$. Modern test…

Methodology · Statistics 2022-12-20 Shubhanshu Shekhar , Ilmun Kim , Aaditya Ramdas

This paper introduces Kernel-based Information Criterion (KIC) for model selection in regression analysis. The novel kernel-based complexity measure in KIC efficiently computes the interdependency between parameters of the model using a…

Machine Learning · Statistics 2014-12-16 Somayeh Danafar , Kenji Fukumizu , Faustino Gomez

In recent years we have been able to gather large amounts of genomic data at a fast rate, creating situations where the number of variables greatly exceeds the number of observations. In these situations, most models that can handle a…

Methodology · Statistics 2025-02-07 Andrea Bratsberg , Abhik Ghosh , Magne Thoresen

As the number of possible predictors generated by high-throughput experiments continues to increase, methods are needed to quickly screen out unimportant covariates. Model-based screening methods have been proposed and theoretically…

Methodology · Statistics 2012-05-31 Sihai D. Zhao , Yi Li

Hilbert-Schmidt independence criterion and distance covariance are methods to describe independence of random variables using either the Kronecker product of positive definite kernels or the Kronecker product of conditionally negative…

Functional Analysis · Mathematics 2022-01-05 Jean Carlo Guella

Model selection is an indispensable part of data analysis dealing very frequently with fitting and prediction purposes. In this paper, we tackle the problem of model selection in a general linear regression where the parameter matrix…

Signal Processing · Electrical Eng. & Systems 2022-09-19 Prakash B. Gohain , Magnus Jansson

Kernel techniques are among the most influential approaches in data science and statistics. Under mild conditions, the reproducing kernel Hilbert space associated to a kernel is capable of encoding the independence of $M\ge 2$ random…

Statistics Theory · Mathematics 2024-10-15 Florian Kalinke , Zoltan Szabo

In this paper, we propose a model-free feature selection method for ultra-high dimensional data with mass features. This is a two phases procedure that we propose to use the fused Kolmogorov filter with the random forest based RFE to remove…

Methodology · Statistics 2023-02-16 Siwei Xia , Yuehan Yang

We propose a nonparametric test of independence, termed optHSIC, between a covariate and a right-censored lifetime. Because the presence of censoring creates a challenge in applying the standard permutation-based testing approaches, we use…

Statistics Theory · Mathematics 2020-11-03 David Rindt , Dino Sejdinovic , David Steinsaltz

We consider the problem of variable screening in ultra-high dimensional generalized linear models (GLMs) of non-polynomial orders. Since the popular SIS approach is extremely unstable in the presence of contamination and noise, we discuss a…

Statistics Theory · Mathematics 2022-11-15 Abhik Ghosh , Erica Ponzi , Torkjel Sandanger , Magne Thoresen

Modeling users' dynamic preferences from historical behaviors lies at the core of modern recommender systems. Due to the diverse nature of user interests, recent advances propose the multi-interest networks to encode historical behaviors…

Information Retrieval · Computer Science 2022-07-19 Zhaocheng Liu , Yingtao Luo , Di Zeng , Qiang Liu , Daqing Chang , Dongying Kong , Zhi Chen

We propose a flexible nonparametric regression method for ultrahigh-dimensional data. As a first step, we propose a fast screening method based on the favored smoothing bandwidth of the marginal local constant regression. Then, an iterative…

Methodology · Statistics 2018-07-30 Yang Feng , Yichao Wu , Leonard Stefanski

Purpose: The scarcity of high-quality curated labeled medical training data remains one of the major limitations in applying artificial intelligence (AI) systems to breast cancer diagnosis. Deep models for mammogram analysis and mass (or…

Computer Vision and Pattern Recognition · Computer Science 2025-04-29 Han Chen , Anne L. Martel

Recent works investigated the generalization properties in deep neural networks (DNNs) by studying the Information Bottleneck in DNNs. However, the mea- surement of the mutual information (MI) is often inaccurate due to the density…

Information Theory · Computer Science 2018-02-16 Denny Wu , Yixiu Zhao , Yao-Hung Hubert Tsai , Makoto Yamada , Ruslan Salakhutdinov

The optimization of high dimensional functions is a key issue in engineering problems but it frequently comes at a cost that is not acceptable since it usually involves a complex and expensive computer code. Engineers often overcome this…

Machine Learning · Statistics 2019-06-18 Adrien Spagnol , Rodolphe Le Riche , Sebastien Da Veiga

Variable selection in high dimensional space has challenged many contemporary statistical problems from many frontiers of scientific disciplines. Recent technology advance has made it possible to collect a huge amount of covariate…

Machine Learning · Statistics 2010-05-20 Jianqing Fan , Yang Feng , Yichao Wu

Refining one's hypotheses in the light of data is a common scientific practice; however, the dependency on the data introduces selection bias and can lead to specious statistical analysis. An approach for addressing this is via conditioning…

Machine Learning · Computer Science 2020-03-03 Jen Ning Lim , Makoto Yamada , Wittawat Jitkrittum , Yoshikazu Terada , Shigeyuki Matsui , Hidetoshi Shimodaira

Identifying important biomarkers that are predictive for cancer patients' prognosis is key in gaining better insights into the biological influences on the disease and has become a critical component of precision medicine. The emergence of…

Methodology · Statistics 2016-03-22 Hyokyoung Grace Hong , Jian Kang , Yi Li
‹ Prev 1 3 4 5 6 7 10 Next ›