Related papers: Maximizing the Bregman divergence from a Bregman f…
The goal of this paper is to further develop an approach to inverse problems with imperfect forward operators that is based on partially ordered spaces. Studying the dual problem yields useful insights into the convergence of the…
We study optimal payoff choice for an expected utility maximizer under the constraint that their payoff is not allowed to deviate ``too much'' from a given benchmark. We solve this problem when the deviation is assessed via a…
Let $\mathcal{A}_1,\ldots,\mathcal{A}_m$ be families of $k$-subsets of an $n$-set. Suppose that one cannot choose pairwise disjoint edges from $s+1$ distinct families. Subject to this condition we investigate the maximum of…
We study many-party correlations quantified in terms of the Umegaki relative entropy (divergence) from a Gibbs family known as a hierarchical model. We derive these quantities from the maximum-entropy principle which was used earlier to…
We present a generalization of the maximal inequalities that upper bound the expectation of the maximum of $n$ jointly distributed random variables. We control the expectation of a randomly selected random variable from $n$ jointly…
The variation distance closure of an exponential family with a convex set of canonical parameters is described, assuming no regularity conditions. The tools are the concepts of convex core of a measure and extension of an exponential…
Logarithmic score and information divergence appear in information theory, statistics, statistical mechanics, and portfolio theory. We demonstrate that all these topics involve some kind of optimization that leads directly to regret…
Projection theorems of divergences enable us to find reverse projection of a divergence on a specific statistical model as a forward projection of the divergence on a different but rather "simpler" statistical model, which, in turn, results…
The ``sample amplification'' problem formalizes the following question: Given $n$ i.i.d. samples drawn from an unknown distribution $P$, when is it possible to produce a larger set of $n+m$ samples which cannot be distinguished from $n+m$…
The growing availability of network data and of scientific interest in distributed systems has led to the rapid development of statistical models of network structure. Typically, however, these are models for the entire network, while the…
We discuss the problem of finding optimal exponents in Diophantine estimates involving one real number and, in some cases where such an exponent is known, present some properties of the corresponding extremal numbers.
We investigate the problem of minimizing Kullback-Leibler divergence between a linear model $Ax$ and a positive vector $b$ in different convex domains (positive orthant, $n$-dimensional box, probability simplex). Our focus is on the SMART…
Bregman divergences play a central role in the design and analysis of a range of machine learning algorithms. This paper explores the use of Bregman divergences to establish reductions between such algorithms and their analyses. We present…
We present here a variational method for maximizing the bandgap in a one-dimensional system where the potential is subject to given constraints. Two specific examples are studied in detail. In the first, we show that if the potential is…
We consider a problem of maximizing the product of the sizes of two uniform cross-$t$-intersecting families of sets. We show that the value of this maximum is at most polynomially larger (in the size of a ground set) than a quantity…
Probabilistic models are often trained by maximum likelihood, which corresponds to minimizing a specific f-divergence between the model and data distribution. In light of recent successes in training Generative Adversarial Networks,…
We formulate the problem of perception in the framework of information theory, and prove that categorical perception is equivalent to the existence of an observable that has the maximum possible information on the target of perception. We…
f-divergence estimation is an important problem in the fields of information theory, machine learning, and statistics. While several divergence estimators exist, relatively few of their convergence rates are known. We derive the MSE…
We propose a novel Bregman descent algorithm for minimizing a convex function that is expressed as the sum of a differentiable part (defined over an open set) and a possibly nonsmooth term. The approach, referred to as the Variable Bregman…
In this paper, we establish the links between the Lehmer and H\"older mean families and maximum weighted likelihood estimator. Considering the regular one-parameter exponential family of probability density functions, we show that the…