English
Related papers

Related papers: Estimating Feature-Label Dependence Using Gini Dis…

200 papers

Recent results in coupled or temporal graphical models offer schemes for estimating the relationship structure between features when the data come from related (but distinct) longitudinal sources. A novel application of these ideas is for…

Machine Learning · Statistics 2017-11-22 Ronak Mehta , Hyunwoo J. Kim , Shulei Wang , Sterling C. Johnson , Ming Yuan , Vikas Singh

Distance-based unsupervised text classification is a method within text classification that leverages the semantic similarity between a label and a text to determine label relevance. This method provides numerous benefits, including fast…

Computation and Language · Computer Science 2025-10-14 Jens Van Nooten , Andriy Kosar , Guy De Pauw , Walter Daelemans

Random feature ridge regression is often analyzed in the high-dimensional regime under the homogeneous sampling model $x_i=\Sigma^{1/2}x_i'$, where the vectors $x_i'$ have iid entries and the same covariance matrix $\Sigma$ is shared by all…

Machine Learning · Statistics 2026-05-19 Issa-Mbenard Dabo , Jérémie Bigot

Given an imperfect predictor, we exploit additional features at test time to improve the predictions made, without retraining and without knowledge of the prediction function. This scenario arises if training labels or data are proprietary,…

Machine Learning · Computer Science 2021-11-05 Kwang In Kim , James Tompkin

A prescription is presented for a new and practical correlation coefficient, $\phi_K$, based on several refinements to Pearson's hypothesis test of independence of two variables. The combined features of $\phi_K$ form an advantage over…

Methodology · Statistics 2019-03-12 M. Baak , R. Koopman , H. Snoek , S. Klous

Distance correlation is a measure of dependence between two paired random vectors or matrices of arbitrary, not necessarily equal, dimensions. Unlike Pearson correlation, the population distance correlation coefficient is zero if and only…

Methodology · Statistics 2025-06-19 Kontemeniotis Nikolaos , Vargiakakis Rafail , Tsagris Michail

An important challenge in statistical analysis lies in controlling the estimation bias when handling the ever-increasing data size and model complexity of modern data settings. In this paper, we propose a reliable estimation and inference…

The extension of bivariate measures of dependence to non-Euclidean spaces is a challenging problem. The non-linear nature of these spaces makes the generalisation of classical measures of linear dependence (such as the covariance) not…

Statistics Theory · Mathematics 2024-10-10 Meshal Abuqrais , Davide Pigoli

The problem of signal detection using sparse, faint information is closely related to a variety of contemporary statistical problems, including the control of false-discovery rate, and classification using very high-dimensional data. Each…

Statistics Theory · Mathematics 2008-12-18 Peter Hall , Jiashun Jin

Mutual Information (MI) is an useful tool for the recognition of mutual dependence berween data sets. Differen methods for the estimation of MI have been developed when both data sets are discrete or when both data sets are continuous. The…

Applications · Statistics 2017-08-30 Miguel A. Ré , Guillermo G. Aguirre Varela

We study the problem of testing $H_0: \xi^\top\beta=t_0$ in high-dimensional sparse linear regression with Gaussian random design and unknown design covariance. The loading vector $\xi$ is arbitrary, and the exact sparsity level $k$ is…

Statistics Theory · Mathematics 2026-05-21 Jie Xie , Dongming Huang

We propose three measures of mutual dependence between multiple random vectors. All the measures are zero if and only if the random vectors are mutually independent. The first measure generalizes distance covariance from pairwise dependence…

Statistics Theory · Mathematics 2018-05-18 Ze Jin , David S. Matteson

Graph neural networks (GNNs) and label propagation represent two interrelated modeling strategies designed to exploit graph structure in tasks such as node property prediction. The former is typically based on stacked message-passing layers…

Machine Learning · Computer Science 2021-10-15 Yangkun Wang , Jiarui Jin , Weinan Zhang , Yongyi Yang , Jiuhai Chen , Quan Gan , Yong Yu , Zheng Zhang , Zengfeng Huang , David Wipf

Recently, graph (network) data is an emerging research area in artificial intelligence, machine learning and statistics. In this work, we are interested in whether node's labels (people's responses) are affected by their neighbor's features…

Methodology · Statistics 2022-10-12 Haixiang Zhang , Yingjun Deng , Alan J. X. Guo , Qing-Hu Hou , Ou Wu

The aim of this thesis is to find a solution to the non-parametric independence problem in separable metric spaces. Suppose we are given finite collection of samples from an i.i.d. sequence of paired random elements, where each marginal has…

Statistics Theory · Mathematics 2017-06-13 Martin Emil Jakobsen

Random features provide a practical framework for large-scale kernel approximation and supervised learning. It has been shown that data-dependent sampling of random features using leverage scores can significantly reduce the number of…

Machine Learning · Computer Science 2019-03-21 Shahin Shahrampour , Soheil Kolouri

We present a study of the dependencies of shear bias on simulation (input) and measured (output) parameters, noise, point-spread function anisotropy, pixel size, and the model bias coming from two different and independent galaxy shape…

Cosmology and Nongalactic Astrophysics · Physics 2020-10-14 Arnau Pujol , Florent Sureau , Jerome Bobin , Frederic Courbin , Marc Gentile , Martin Kilbinger

During the training process, deep neural networks implicitly learn to represent the input data samples through a hierarchy of features, where the size of the hierarchy is determined by the number of layers. In this paper, we focus on…

Computer Vision and Pattern Recognition · Computer Science 2022-04-08 Florinel-Alin Croitoru , Diana-Nicoleta Grigore , Radu Tudor Ionescu

Datasets may contain observations with multiple labels. If the labels are not mutually exclusive, and if the labels vary greatly in frequency, obtaining a sample that includes sufficient observations with scarcer labels to make inferences…

Machine Learning · Computer Science 2026-05-27 Simon Chung , Colby J. Vorland , Donna L. Maney , Andrew W. Brown

Identifying how dependence relationships vary across different conditions plays a significant role in many scientific investigations. For example, it is important for the comparison of biological systems to see if relationships between…

Methodology · Statistics 2023-07-31 Hoseung Song , Michael C. Wu