Related papers: Measuring Dependencies of Order Statistics: An Inf…
A class of probability distributions is characterized via equalities in law between two order statistics shifted by independent exponential variables. An explicit formula for the quintile function of the identified family of distributions…
The general relationship between an arbitrary frequency distribution and the expectation value of the frequency distributions of its samples is esablished. A set of combinations of expectation values whose value does not in general depend…
Several tasks in information retrieval (IR) rely on assumptions regarding the distribution of some property (such as term frequency) in the data being processed. This thesis argues that such distributional assumptions can lead to incorrect…
This study focuses on statistical inference for compound models of the form $X=\xi_1+\ldots+\xi_N$, where $N$ is a random variable denoting the count of summands, which are independent and identically distributed (i.i.d.) random variables…
In this article, we study tests of independence for data with arbitrary distributions in the non-serial case, i.e., for independent and identically distributed random vectors, as well as in the serial case, i.e., for time series. These…
We give the distribution of $M_n$, the maximum of a sequence of $n$ observations from a moving average of order 1. Solutions are first given in terms of repeated integrals and then for the case where the underlying independent random…
We discuss the connection between information and copula theories by showing that a copula can be employed to decompose the information content of a multivariate distribution into marginal and dependence components, with the latter…
Let $X_1, X_2,\ldots, X_n$ be $n$ independent and identically distributed random variables, here $n \geq 2.$ Let $X_{(1)}, X_{(2)}, \ldots, X_{(n)}$ be the order statistics of $X_1, X_2,..., X_n.$ In this note we proved that: (I) If $X_1,…
Real-life data are often non-IID due to complex distributions and interactions, and the sensitivity to the distribution of samples can differ among learning models. Accordingly, a key question for any supervised or unsupervised model is…
Explicit finite-sample statistical guarantees on model performance are an important ingredient in responsible machine learning. Previous work has focused mainly on bounding either the expected loss of a predictor or the probability that an…
Statistical learning relies upon data sampled from a distribution, and we usually do not care what actually generated it in the first place. From the point of view of causal modeling, the structure of each distribution is induced by…
This thesis details a class of partial orders on the space of probability distributions and the space of density operators which capture the idea of information content. Some links to domain theory and computational linguistics are also…
Measuring dependence between random variables is a fundamental problem in Statistics, with applications across diverse fields. While classical measures such as Pearson's correlation have been widely used for over a century, they have…
This paper proposes a new statistic to test independence between two high dimensional random vectors ${\mathbf{X}}:p_1\times1$ and ${\mathbf{Y}}:p_2\times1$. The proposed statistic is based on the sum of regularized sample canonical…
Testing the independence between random vectors is a fundamental problem in statistics. Distance correlation, a recently popular dependence measure, is universally consistent for testing independence against all distributions with finite…
We propose a new nonparametric test for the supposition of independence between two continuous random variables. The test is based on the size of the longest increasing subsequence of a random permutation. We identified the independence…
Consider independent observations $(X_1,R_1)$, $(X_2,R_2)$, \ldots, $(X_n,R_n)$ with random or fixed ranks $R_i \in \{1,2,\ldots,k\}$, while conditional on $R_i = r$, the random variable $X_i$ has the same distribution as the $r$-th order…
Dependency networks (Heckerman et al., 2000) provide a flexible framework for modeling complex systems with many variables by combining independently learned local conditional distributions through pseudo-Gibbs sampling. Despite their…
Let $X_{\lambda _{1}},X_{\lambda _{2}},\ldots ,X_{\lambda _{n}}$ be independent nonnegative random variables with $X_{\lambda _{i}}\sim F(\lambda _{i}t)$, $i=1,\ldots ,n$, where $\lambda _{i}>0$, $i=1,\ldots ,n$ and $F$ is an absolutely…
How can the information that a set ${X_{1},...,X_{n}}$ of random variables contains about another random variable $S$ be decomposed? To what extent do different subgroups provide the same, i.e. shared or redundant, information, carry unique…