English
Related papers

Related papers: Theoretical Foundations of Equitability and the Ma…

200 papers

A measure of dependence is said to be equitable if it gives similar scores to equally noisy relationships of different types. Equitability is important in data exploration when the goal is to identify a relatively small set of strongest…

Machine Learning · Computer Science 2013-08-16 David Reshef , Yakir Reshef , Michael Mitzenmacher , Pardis Sabeti

Reshef et al. recently proposed a new statistical measure, the "maximal information coefficient" (MIC), for quantifying arbitrary dependencies between pairs of stochastic quantities. MIC is based on mutual information, a fundamental…

Quantitative Methods · Quantitative Biology 2015-06-12 Justin B. Kinney , Gurinder S. Atwal

Given a high-dimensional data set we often wish to find the strongest relationships within it. A common strategy is to evaluate a measure of dependence on every variable pair and retain the highest-scoring pairs for follow-up. This strategy…

The Maximal Information Coefficient (MIC) of Reshef et al. (Science, 2011) is a statistic for measuring dependence between variable pairs in large datasets. In this note, we prove that MIC is a consistent estimator of the corresponding…

Methodology · Statistics 2021-07-09 John Lazarsfeld , Aaron Johnson

Reshef & Reshef recently published a paper in which they present a method called the Maximal Information Coefficient (MIC) that can detect all forms of statistical dependence between pairs of variables as sample size goes to infinity. While…

Machine Learning · Statistics 2013-08-28 Alexander Luedtke , Linh Tran

The maximal information coefficient (MIC), which measures the amount of dependence between two variables, is able to detect both linear and non-linear associations. However, computational cost grows rapidly as a function of the dataset…

Information Theory · Computer Science 2015-08-18 Ali Mousavi , Richard G. Baraniuk

In exploratory data analysis, we are often interested in identifying promising pairwise associations for further analysis while filtering out weaker, less interesting ones. This can be accomplished by computing a measure of dependence on…

Methodology · Statistics 2018-03-28 David N. Reshef , Yakir A. Reshef , Pardis C. Sabeti , Michael M. Mitzenmacher

For analysis of a high-dimensional dataset, a common approach is to test a null hypothesis of statistical independence on all variable pairs using a non-parametric measure of dependence. However, because this approach attempts to identify…

Statistics Theory · Mathematics 2015-05-14 Yakir A. Reshef , David N. Reshef , Pardis C. Sabeti , Michael M. Mitzenmacher

In Science, Reshef et al. (2011) proposed the concept of equitability for measures of dependence between two random variables. To this end, they proposed a novel measure, the maximal information coefficient (MIC). Recently a PNAS paper…

Methodology · Statistics 2023-04-17 A. Adam Ding , Yi Li

The Maximal Information Coefficient (MIC) is a powerful statistic to identify dependencies between variables. However, it may be applied to sensitive data, and publishing it could leak private information. As a solution, we present…

Cryptography and Security · Computer Science 2022-06-23 John Lazarsfeld , Aaron Johnson , Emmanuel Adeniran

Motivation: Clustering is a frequently used concept in variety of bioinformatical applications. We present a new method for hierarchical clustering of data called mutual information clustering (MIC) algorithm. It uses mutual information…

Quantitative Methods · Quantitative Biology 2007-05-23 Alexander Kraskov , Harald Stögbauer , Ralph G. Andrzejak , Peter Grassberger

Clustering is a concept used in a huge variety of applications. We review a conceptually very simple algorithm for hierarchical clustering called in the following the {\it mutual information clustering} (MIC) algorithm. It uses mutual…

Quantitative Methods · Quantitative Biology 2008-09-10 Alexander Kraskov , Peter Grassberger

A computational framework utilizes the traditional similarity measures for mining the significant relationships in biological annotations is recently proposed by Tatiana V. Karpinets et al. [2]. In this paper, an improved approximation…

Databases · Computer Science 2015-07-21 Shuliang Wang , Yiping Zhao

When two variables are related by a known function, the coefficient of determination (denoted $R^2$) measures the proportion of the total variance in the observations that is explained by that function. This quantifies the strength of the…

Applications · Statistics 2013-03-11 Ben Murrell , Daniel Murrell , Hugh Murrell

Mutual Information (MI) is a powerful statistical measure that quantifies shared information between random variables, particularly valuable in high-dimensional data analysis across fields like genomics, natural language processing, and…

Machine Learning · Computer Science 2024-12-02 Andre O. Falcao

Total correlation (TC) is a fundamental concept in information theory that measures statistical dependency among multiple random variables. Recently, TC has shown noticeable effectiveness as a regularizer in many learning tasks, where the…

Information Theory · Computer Science 2023-02-23 Ke Bai , Pengyu Cheng , Weituo Hao , Ricardo Henao , Lawrence Carin

We introduce a new information theoretic measure that we call Public Information Complexity (PIC), as a tool for the study of multi-party computation protocols, and of quantities such as their communication complexity, or the amount of…

Computational Complexity · Computer Science 2018-12-18 Iordanis Kerenidis , Adi Rosén , Florent Urrutia

The proposal of Reshef et al. (2011) is an interesting new approach for discovering non-linear dependencies among pairs of measurements in exploratory data mining. However, it has a potentially serious drawback. The authors laud the fact…

Methodology · Statistics 2014-01-30 Noah Simon , Robert Tibshirani

Relational data augmentation is a powerful technique for enhancing data analytics and improving machine learning models by incorporating columns from external datasets. However, it is challenging to efficiently discover relevant external…

Databases · Computer Science 2025-03-06 Aécio Santos , Flip Korn , Juliana Freire

Maximum mutual information (MMI) is a model selection criterion used for hidden Markov model (HMM) parameter estimation that was developed more than twenty years ago as a discriminative alternative to the maximum likelihood criterion for…

Computation and Language · Computer Science 2010-02-04 Steven Wegmann
‹ Prev 1 2 3 10 Next ›