English
Related papers

Related papers: Sharp Inequalities between Total Variation and Hel…

200 papers

We provide a general bound on the Wasserstein distance between two arbitrary distributions of sequences of Bernoulli random variables. The bound is in terms of a mixing quantity for the Glauber dynamics of one of the sequences, and a simple…

Probability · Mathematics 2018-10-12 Gesine Reinert , Nathan Ross

Generalization error bounds are essential to understanding machine learning algorithms. This paper presents novel expected generalization error upper bounds based on the average joint distribution between the output hypothesis and each…

Information Theory · Computer Science 2022-02-25 Gholamali Aminian , Yuheng Bu , Gregory Wornell , Miguel Rodrigues

We deal with stochastic differential equations with jumps. In order to obtain an accurate approximation scheme, it is usual to replace the "small jumps" by a Brownian motion. In this paper, we prove that for every fixed time $t$, the…

Probability · Mathematics 2023-10-17 Vlad Bally , Yifeng Qin

The Jeffreys divergence is a renown symmetrization of the oriented Kullback-Leibler divergence broadly used in information sciences. Since the Jeffreys divergence between Gaussian mixture models is not available in closed-form, various…

Information Theory · Computer Science 2021-11-24 Frank Nielsen

In this paper we provide explicit upper bounds on some distances between the (law of the) output of a random Gaussian NN and (the law of) a random Gaussian vector. Our results concern both shallow random Gaussian neural networks with…

This article studies the infinite-width limit of deep feedforward neural networks whose weights are dependent, and modelled via a mixture of Gaussian distributions. Each hidden node of the network is assigned a nonnegative random variable…

Machine Learning · Statistics 2025-02-06 Hoil Lee , Fadhel Ayed , Paul Jung , Juho Lee , Hongseok Yang , François Caron

We consider the initial situation where a dataset has been over-partitioned into $k$ clusters and seek a domain independent way to merge those initial clusters. We identify the total variation distance (TVD) as suitable for this goal. By…

Machine Learning · Computer Science 2019-12-10 Christian Reiser , Jörg Schlötterer , Michael Granitzer

The Tully-Fisher relation is a vital distance indicator, but its precise inference is challenged by selection bias, statistical bias, and uncertain inclination corrections. This study presents a Bayesian framework that simultaneously…

Methodology · Statistics 2025-09-09 Hai Fu

The likelihood function of a finite mixture model is a non-convex function with multiple local maxima and commonly used iterative algorithms such as EM will converge to different solutions depending on initial conditions. In this paper we…

Machine Learning · Computer Science 2016-08-19 Elad Mezuman , Yair Weiss

These expository notes introduce the Hellinger distance on the set of all measures and the induced Fisher-Rao distances for subsets of measures, such as probability measures or Gaussian measures. The historical background is highlighted and…

Mathematical Physics · Physics 2025-10-06 Alexander Mielke

In this paper, we develop local expansions for the ratio of the centered matrix-variate $T$ density to the centered matrix-variate normal density with the same covariances. The approximations are used to derive upper bounds on several…

Statistics Theory · Mathematics 2022-11-18 Frédéric Ouimet

Mixtures of Gaussian (or normal) distributions arise in a variety of application areas. Many heuristics have been proposed for the task of finding the component Gaussians given samples from the mixture, such as the EM algorithm, a…

Probability · Mathematics 2007-05-23 Sanjeev Arora , Ravi Kannan

Recently, a Wasserstein-type distance for Gaussian mixture models has been proposed. However, that framework can only be generalized to identifiable mixtures of general elliptically contoured distributions whose components come from the…

Optimization and Control · Mathematics 2025-03-19 Keyu Chen , Zetian Wang , Yunxin Zhang

We deal with stochastic differential equations with jumps. In order to obtain an accurate approximation scheme, it is usual to replace the "small jumps" by a Brownian motion. In this paper, we prove that for every fixed time $t$, the…

Probability · Mathematics 2022-12-15 Vlad Bally , Yifeng Qin

In a landmark result, Chen et al. (2018) showed that multivariate medians induced by halfspace depth attain the minimax optimal convergence rate under Huber contamination and elliptical symmetry, for both location and scatter estimation. We…

Statistics Theory · Mathematics 2025-12-19 Filip Bočinec , Stanislav Nagy

We prove that $\tilde{\Theta}(k d^2 / \varepsilon^2)$ samples are necessary and sufficient for learning a mixture of $k$ Gaussians in $\mathbb{R}^d$, up to error $\varepsilon$ in total variation distance. This improves both the known upper…

Machine Learning · Computer Science 2020-07-23 Hassan Ashtiani , Shai Ben-David , Nick Harvey , Christopher Liaw , Abbas Mehrabian , Yaniv Plan

We study the problem of distributed multi-view representation learning. In this problem, $K$ agents observe each one distinct, possibly statistically correlated, view and independently extracts from it a suitable representation in a manner…

Machine Learning · Statistics 2025-04-28 Milad Sefidgaran , Abdellatif Zaidi , Piotr Krasnowski

Tight bounds for several symmetric divergence measures are derived in terms of the total variation distance. It is shown that each of these bounds is attained by a pair of 2 or 3-element probability distributions. An application of these…

Information Theory · Computer Science 2016-11-17 Igal Sason

We prove results on the decidability and complexity of computing the total variation distance (equivalently, the $L_1$-distance) of hidden Markov models (equivalently, labelled Markov chains). This distance measures the difference between…

Formal Languages and Automata Theory · Computer Science 2018-04-18 Stefan Kiefer

Distance measures between graphs are important primitives for a variety of learning tasks. In this work, we describe an unsupervised, optimal transport based approach to define a distance between graphs. Our idea is to derive…

Computational Engineering, Finance, and Science · Computer Science 2024-04-11 Michael Scholkemper , Damin Kühn , Gerion Nabbefeld , Simon Musall , Björn Kampa , Michael T. Schaub