Related papers: On Suspicious Coincidences and Pointwise Mutual In…
Given the increasing popularity of algorithms for overlapping clustering, in particular in social network analysis, quantitative measures are needed to measure the accuracy of a method. Given a set of true clusters, and the set of clusters…
Composite likelihood inference has gained much popularity thanks to its computational manageability and its theoretical properties. Unfortunately, performing composite likelihood ratio tests is inconvenient because of their awkward…
During a speculative episode the price of an item jumps from an initial level p_1 to a peak level p_2 before more or less returning to level p_1. The ratio p_2/p_1 is referred to as the amplitude A of the peak. This paper shows that for a…
Let $S$ and $\tilde S$ be two independent and identically distributed random variables, which we interpret as the signal, and let $P_1$ and $P_2$ be two communication channels. We can choose between two measurement scenarios: either we…
Testing hypothesis of independence between two random elements on a joint alphabet is a fundamental exercise in statistics. Pearson's chi-squared test is an effective test for such a situation when the contingency table is relatively small.…
For discrete random variables X_1,..., X_n we construct an n by n matrix. In the (i,j) entry we put the mutual information I(X_i;X_j) between X_i and X_j. In particular, in the (i,i) entry we put the entropy H(X_i)=I(X_i;X_i) of X_i. This…
I present several new relations between mutual information (MI) and statistical estimation error for a system that can be regarded simultaneously as a communication channel and as an estimator of an input parameter. I first derive a…
The amount of mutual information contained in time series of two elements gives a measure of how well their activities are coordinated. In a large, complex network of interacting elements, such as a genetic regulatory network within a cell,…
One can often encounter claims that classical (Kolmogorovian) probability theory cannot handle, or even is contradicted by, certain empirical findings or substantive theories. This note joins several previous attempts to explain that these…
Responsible indicators are crucial for research assessment and monitoring. Transparency and accuracy of indicators are required to make research assessment fair and ensure reproducibility. However, sometimes it is difficult to conduct or…
We propose the concept of mutual information for particle pair (MIPP) in curved spacetime, and show that MIPP has potential to be a proper chaos indicator. We tested this method in the Schwarzschild and Kerr spacetime and compared it with…
In data science, it is often required to estimate dependencies between different data sources. These dependencies are typically calculated using Pearson's correlation, distance correlation, and/or mutual information. However, none of these…
Ranking and comparing items is crucial for collecting information about preferences in many areas, from marketing to politics. The Mallows rank model is among the most successful approaches to analyse rank data, but its computational…
Learning from positive and unlabeled data (PU learning) is actively researched machine learning task. The goal is to train a binary classification model based on a training dataset containing part of positives which are labeled, and…
Quantifying cooperation or synergy among random variables in predicting a single target random variable is an important problem in many complex systems. We review three prior information-theoretic measures of synergy and introduce a novel…
This paper considers the model of an arbitrary distributed signal x observed through an added independent white Gaussian noise w, y=x+w. New relations between the minimal mean square error of the non-causal estimator and the likelihood…
If $S$ and $T$ are infinite sequences over a finite alphabet, then the lower and upper mutual dimensions $mdim(S:T)$ and $Mdim(S:T)$ are the upper and lower densities of the algorithmic information that is shared by $S$ and $T$. In this…
P values or risk ratios from multiple, independent studies, observational or randomized, can be computationally combined to provide an overall assessment of a research question in meta-analysis. There is a need to examine the reliability of…
Shannon defined the mutual information between two variables. We illustrate why the true mutual information between a variable and the predictions made by a prediction algorithm is not a suitable measure of prediction quality, but the…
Bayesian neural networks have successfully designed and optimized a robust neural network model in many application problems, including uncertainty quantification. However, with its recent success, information-theoretic understanding about…