Related papers: Minimum HGR Correlation Principle: From Marginals …
Consider the binary classification problem of predicting a target variable $Y$ from a discrete feature vector $X = (X_1,...,X_d)$. When the probability distribution $\mathbb{P}(X,Y)$ is known, the optimal classifier, leading to the minimum…
The Hirschfeld-Gebelein-R\'enyi (HGR) correlation coefficient is an extension of Pearson's correlation that is not limited to linear correlations, with potential applications in algorithmic fairness, scientific analysis, and causal…
For independent random variables $(X_i)_{1\leq i\leq n}$, we consider the maximal correlation coefficient $R=R(\min_{i:1\leq i\leq m}X_i,\min_{j:\ell+1\leq j\leq n}X_j)$. If $X_1,X_2,\ldots,X_n$ are identically distributed with the same…
Given two discrete random variables $X$ and $Y$, with probability distributions ${\bf p} =(p_1, \ldots , p_n)$ and ${\bf q}=(q_1, \ldots , q_m)$, respectively, denote by ${\cal C}({\bf p}, {\bf q})$ the set of all couplings of ${\bf p}$ and…
Consider a bivariate Geometric random variable where the first component has parameter $p_1$ and the second parameter $p_2$. It is not possible to make the correlation between the marginals equal to -1. Here the properties of this minimum…
The maximal (or Hilbertian) correlation coefficient between two random variables X and Y, denoted by \{X:Y\}, is the supremum of the |Corr(f(X),g(Y))| for real measurable functions f, g, where "Corr" denotes Pearson's correlation…
The joint distribution $P(X,Y)$ cannot be determined from its marginals $P(X)$ and $P(Y)$ alone; one also needs one of the conditionals $P(X|Y)$ or $P(Y|X)$. But is there a best guess, given only the marginals? Here we answer this question…
Consider the problem of drawing random variates $(X_1,\ldots,X_n)$ from a distribution where the marginal of each $X_i$ is specified, as well as the correlation between every pair $X_i$ and $X_j$. For given marginals, the…
Given two discrete random variables $X$ and $Y,$ with probability distributions ${\bf p}=(p_1, \ldots , p_n)$ and ${\bf q}=(q_1, \ldots , q_m)$, respectively, denote by ${\cal C}({\bf p}, {\bf q})$ the set of all couplings of ${\bf p}$ and…
We consider a distributed logistic regression problem where labeled data pairs $(X_i,Y_i)\in \mathbb{R}^d\times\{-1,1\}$ for $i=1,\ldots,n$ are distributed across multiple machines in a network and must be communicated to a centralized…
We propose a lower bound on the log marginal likelihood of Gaussian process regression models that can be computed without matrix factorisation of the full kernel matrix. We show that approximate maximum likelihood learning of model…
We correct claims about lower bounds on mutual information (MI) between real-valued random variables made in A. Kraskov {\it et al.}, Phys. Rev. E {\bf 69}, 066138 (2004). We show that non-trivial lower bounds on MI in terms of linear…
Hybrid Bayesian Networks (HBNs), which contain both discrete and continuous variables, arise naturally in many application areas (e.g., image understanding, data fusion, medical diagnosis, fraud detection). This paper concerns inference in…
We construct optimal low-rank approximations for the Gaussian posterior distribution in linear Gaussian inverse problems with possibly infinite-dimensional separable Hilbert parameter spaces and finite-dimensional data spaces. We first…
The Hirschfeld-Gebelein-R\'{e}nyi (HGR) maximal correlation and the corresponding functions have been shown useful in many machine learning scenarios. In this paper, we study the sample complexity of estimating the HGR maximal correlation…
Many types of bounded data defined on the unit interval arise naturally as ratios of the form $X/(X + Y)$. In the existing literature, the main statistical models proposed for this type of bounded data typically based on the assumption that…
Graph cuts are among the most prominent tools for clustering and classification analysis. While intensively studied from geometric and algorithmic perspectives, graph cut-based statistical inference still remains elusive to a certain…
Many inference problems involving questions of optimality ask for the maximum or the minimum of a finite set of unknown quantities. This technical report derives the first two posterior moments of the maximum of two correlated Gaussian…
George R. Terrell (1983, {Ann. Probab., vol. 11(3), pp. 823--826) showed that the Pearson coefficient of correlation of an ordered pair from a random sample of size two is at most one-half, and the equality is attained only for rectangular…
Let $X=(x_{ij})\in\mathbb{R}^{N\times n}$ be a rectangular random matrix with i.i.d. entries (we assume $N/n\to\mathbf{a}>1$), and denote by $\sigma_{min}(X)$ its smallest singular value. When entries have mean zero and unit second moment,…