Related papers: Hyper Normalisation and Conditioning for Discrete …
Statistical thermodynamics delivers the probability distribution of the equilibrium state of matter through the constrained maximization of a special functional, entropy. Its elegance and enormous success have led to numerous attempts to…
We generalize McDiarmid's inequality for functions with bounded differences on a high probability set, using an extension argument. Those functions concentrate around their conditional expectations. We further extend the results to…
Quantization for a Borel probability measure refers to the idea of estimating a given probability by a discrete probability with support containing a finite number of elements. If in the quantization some of the elements in the support are…
In probability theory, there is a tendency to treat one random variable with a given distribution as being just as good as any other. By and large this is fine because probability is (mostly) concerned with distributional properties of…
The importance of properly taking into account the factorization scheme dependence of parton distribution functions is emphasized. A serious error in the usual handling of this topic is pointed out and the correct procedure for transforming…
After a short review of the historical milestones on normal numbers, we introduce the Borel numbers as the reals admitting a probability function on their different bases representations. In this setting, we provide two probabilistic…
General classes of bivariate distributions are well studied in literature. Most of these classes are proposed via a copula formulation or extensions of some characterisation properties in the univariate case. In Kundu(2022) we see one such…
An equivalent condition for the product of elements of an independent random sample on a compact algebraic group converging in distribution to some random variable as the sample size increases is obtained. Namely, a limit distribution…
A central limit theorem is established for a sum of random variables belonging to a sequence of random fields. The fields are assumed to have zero mean conditional on the past history and to satisfy certain conditional $\alpha$-mixing…
Normalization techniques have only recently begun to be exploited in supervised learning tasks. Batch normalization exploits mini-batch statistics to normalize the activations. This was shown to speed up training and result in better…
On-line learning of probability distributions is analyzed from the field theoretical point of view. We can obtain an optimal on-line learning algorithm, since renormalization group enables us to control the number of degrees of freedom of a…
Recent literature in the last Maximum Entropy workshop introduced an analogy between cumulative probability distributions and normalized utility functions. Based on this analogy, a utility density function can de defined as the derivative…
The generation of comprehensible explanations is an essential feature of modern artificial intelligence systems. In this work, we consider probabilistic logic programming, an extension of logic programming which can be useful to model…
The goal of regression and classification methods in supervised learning is to minimize the empirical risk, that is, the expectation of some loss function quantifying the prediction error under the empirical distribution. When facing scarce…
We present a general theorem on the structure of bivariate generating functions which gives sufficient conditions such that the limiting probability distribution is a half-normal distribution. If $X$ is a normally distributed random…
Traditional normalization techniques (e.g., Batch Normalization and Instance Normalization) generally and simplistically assume that training and test data follow the same distribution. As distribution shifts are inevitable in real-world…
Theory refinement is the task of updating a domain theory in the light of new cases, to be done automatically or with some expert assistance. The problem of theory refinement under uncertainty is reviewed here in the context of Bayesian…
The problem of minimizing convex functionals of probability distributions is solved under the assumption that the density of every distribution is bounded from above and below. A system of sufficient and necessary first-order optimality…
This PhD thesis presents a distributional view of optimization in place of a worst-case perspective. We motivate this view with an investigation of the failure point of classical optimization. Subsequently we consider the optimization of a…
Bayesian inference gets its name from *Bayes's theorem*, expressing posterior probabilities for hypotheses about a data generating process as the (normalized) product of prior probabilities and a likelihood function. But Bayesian inference…