Related papers: Sum decomposition of divergence into three diverge…
Measuring the distance between data points is fundamental to many statistical techniques, such as dimension reduction or clustering algorithms. However, improvements in data collection technologies has led to a growing versatility of…
In this paper, using some aspects of convex functions, we refine discrete Jensen's inequality via weight functions. Then, using these results, we give some applications in different abstract spaces and obtain some new interesting…
Deep embeddings answer one simple question: How similar are two images? Learning these embeddings is the bedrock of verification, zero-shot learning, and visual search. The most prominent approaches optimize a deep convolutional network…
Dedekind sums have applications in quite a number of fields of mathematics. Therefore, their distribution has found considerable interest. This article gives a survey of several aspects of the distribution of these sums. In particular, it…
Let $i(\infty,k)$ be the limiting proportion, as $n \rightarrow \infty$, of permutations in the symmetric group of degree $n$ that fix a $k$-set. We give an algorithm for computing $i(\infty,k)$ and state the values of $i(\infty,k)$ for $k…
Estimating the ratio of two probability densities from a finite number of observations is a central machine learning problem. A common approach is to construct estimators using binary classifiers that distinguish observations from the two…
Variational inequalities play a key role in machine learning research, such as generative adversarial networks, reinforcement learning, adversarial training, and generative models. This paper is devoted to the constrained variational…
We investigate two classes of transformations of cosine similarity and Pearson and Spearman correlations into metric distances, utilising the simple tool of metric-preserving functions. The first class puts anti-correlated objects maximally…
In this paper, sums represented in (3) are studied. The expressions are derived in terms of Bessel functions of the first and second kinds and their integrals. Further, we point out the integrals can be written as a Meijer G function.
We investigate a result on convergence of double sequences of numbers and how it extends to measurable functions.
The study of finite approximations of probability measures has a long history. In (Xu and Berger, 2017), the authors focus on constrained finite approximations and, in particular, uniform ones in dimension $d=1$. The present paper gives an…
The estimation of an f-divergence between two probability distributions based on samples is a fundamental problem in statistics and machine learning. Most works study this problem under very weak assumptions, in which case it is provably…
A loss function measures the discrepancy between the true values (observations) and their estimated fits, for a given instance of data. A loss function is said to be proper (unbiased, Fisher consistent) if the fits are defined over a unit…
In knowledge graph embedding, the theoretical relationship between the softmax cross-entropy and negative sampling loss functions has not been investigated. This makes it difficult to fairly compare the results of the two different loss…
The geometric Jensen--Shannon divergence (G-JSD) gained popularity in machine learning and information sciences thanks to its closed-form expression between Gaussian distributions. In this work, we introduce an alternative definition of the…
Contragenic functions are defined to be reduced-quaternion-valued harmonic functions which are orthogonal to all monogenic and antimonogenic functions in the $L^2$ norm of a given domain. The parallelism between the spaces of contragenic…
This expository paper presents elementary proofs of four basic results concerning derivatives of quasi-convex functions. They are combined into a fifth theorem which is simple to apply and adequate in many cases. Along the way we establish…
Diversity is a central concept in many fields. Despite its importance, there is no unified methodological framework to measure diversity and its three components of variety, balance and disparity. Current approaches take into account…
The new ingredient of this paper is that we consider infinitely dimensional classes of functions and instead of the relative error setting, which was used in previous papers on norm discretization, we consider the absolute error setting. We…
We give a new interpretation of the derangement numbers d_n as the sum of the values of the largest fixed points of all non-derangements of length n-1. We also show that the analogous sum for the smallest fixed points equals the number of…