English
Related papers

Related papers: Estimating Feature-Label Dependence Using Gini Dis…

200 papers

Graphical models are widely used in diverse application domains to model the conditional dependencies amongst a collection of random variables. In this paper, we consider settings where the graph structure is covariate-dependent, and…

Machine Learning · Statistics 2025-04-24 Jiahe Lin , Yikai Zhang , George Michailidis

We propose the conditional predictive impact (CPI), a consistent and unbiased estimator of the association between one or several features and a given outcome, conditional on a reduced feature set. Building on the knockoff framework of…

Methodology · Statistics 2021-05-14 David S. Watson , Marvin N. Wright

An open scientific challenge is how to classify events with reliable measures of uncertainty, when we have a mechanistic model of the data-generating process but the distribution over both labels and latent nuisance parameters is different…

Machine Learning · Statistics 2024-07-02 Luca Masserano , Alex Shen , Michele Doro , Tommaso Dorigo , Rafael Izbicki , Ann B. Lee

Pearson's Chi-squared test, though widely used for detecting association between categorical variables, exhibits low statistical power in large sparse contingency tables. To address this limitation, two novel permutation tests have been…

Methodology · Statistics 2024-03-27 Qingyang Zhang

Simple correlation coefficients between two variables have been generalized to measure association between two matrices in many ways. Coefficients such as the RV coefficient, the distance covariance (dCov) coefficient and kernel based…

Methodology · Statistics 2014-08-19 Julie Josse , Susan Holmes

We study the problem of learning feature representations from a pair of random variables, where we focus on the representations that are induced by their dependence. We provide sufficient and necessary conditions for such dependence induced…

Machine Learning · Computer Science 2024-11-26 Xiangxiang Xu , Lizhong Zheng

In this paper, we propose two new flexible Gini indices (extended lower and upper) defined via differences between the $i$-th observation, the smallest order statistic, and the largest order statistic, for any $1 \leqslant i \leqslant m$.…

Methodology · Statistics 2025-06-03 Roberto Vila , Helton Saulo

Understanding and developing a correlation measure that can detect general dependencies is not only imperative to statistics and machine learning, but also crucial to general scientific discovery in the big data age. In this paper, we…

Machine Learning · Statistics 2024-06-27 Cencheng Shen , Carey E. Priebe , Joshua T. Vogelstein

Distance covariance and distance correlation are scalar coefficients that characterize independence of random vectors in arbitrary dimension. Properties, extensions, and applications of distance correlation have been discussed in the recent…

Methodology · Statistics 2014-07-10 Gabor J. Szekely , Maria L. Rizzo

Genome-wide association studies (GWAS) have identified thousands of genetic variants associated with complex traits, and some variants are shown to be associated with multiple complex traits. Genetic covariance between two traits is defined…

Methodology · Statistics 2023-10-06 Jianqiao Wang , Sai Li , Hongzhe Li

The K-sample testing problem involves determining whether K groups of data points are each drawn from the same distribution. Analysis of variance is arguably the most classical method to test mean differences, along with several recent…

Machine Learning · Statistics 2024-10-04 Sambit Panda , Cencheng Shen , Ronan Perry , Jelle Zorn , Antoine Lutz , Carey E. Priebe , Joshua T. Vogelstein

We provide a unified framework for independence and mean independence tests based on the Hilbert-Schmidt independence criterion, extending some previous results in the literature to hold in general topological spaces. We also present a…

Methodology · Statistics 2026-05-01 Daniel Diz-Castro , Manuel Febrero-Bande , Wenceslao González-Manteiga

Gaussian mixture models are widely used to model data generated from multiple latent sources. Despite its popularity, most theoretical research assumes that the labels are either independent and identically distributed, or follows a Markov…

Statistics Theory · Mathematics 2025-10-09 Seunghyun Lee , Rajarshi Mukherjee , Sumit Mukherjee

In a typical supervised machine learning setting, the predictions on all test instances are based on a common subset of features discovered during model training. However, using a different subset of features that is most informative for…

Machine Learning · Computer Science 2021-06-10 Yasitha Warahena Liyanage , Daphney-Stavroula Zois , Charalampos Chelmis

The covariance of two random variables measures the average joint deviations from their respective means. We generalise this well-known measure by replacing the means with other statistical functionals such as quantiles, expectiles, or…

Methodology · Statistics 2023-09-22 Tobias Fissler , Marc-Oliver Pohle

This paper introduces enhancements to the K-means and K-nearest neighbors (KNN) algorithms based on the concept of Gini prametric spaces, instead of traditional metric spaces. Unlike standard distance metrics, Gini prametrics incorporate…

Machine Learning · Computer Science 2025-08-27 Cassandra Mussard , Arthur Charpentier , Stéphane Mussard

Many relations of scientific interest are nonlinear, and even in linear systems distributions are often non-Gaussian, for example in fMRI BOLD data. A class of search procedures for causal relations in high dimensional data relies on sample…

Artificial Intelligence · Computer Science 2014-01-30 Joseph D. Ramsey

This paper introduces a new method for testing the statistical significance of estimated parameters in predictive regressions. The approach features a new family of test statistics that are robust to the degree of persistence of the…

Econometrics · Economics 2025-02-04 Jean-Yves Pitarakis

Measuring dependence between two events, or equivalently between two binary random variables, amounts to expressing the dependence structure inherent in a $2\times 2$ contingency table in a real number between $-1$ and $1$. Countless such…

Methodology · Statistics 2025-11-13 Marc-Oliver Pohle , Timo Dimitriadis , Jan-Lukas Wermuth

Deciphering the associations between network connectivity and nodal attributes is one of the core problems in network science. The dependency structure and high-dimensionality of networks pose unique challenges to traditional dependency…

Methodology · Statistics 2024-06-27 Youjin Lee , Cencheng Shen , Carey E. Priebe , Joshua T. Vogelstein
‹ Prev 1 4 5 6 7 8 10 Next ›