English
Related papers

Related papers: Sliced Distribution Matching based on Cumulative D…

200 papers

There are many applications that benefit from computing the exact divergence between 2 discrete probability measures, including machine learning. Unfortunately, in the absence of any assumptions on the structure or independencies within…

Machine Learning · Computer Science 2023-10-16 Loong Kuan Lee , Nico Piatkowski , François Petitjean , Geoffrey I. Webb

We consider the problem of uniformity testing of Lipschitz continuous distributions with bounded support. The alternative hypothesis is a composite set of Lipschitz continuous distributions that are at least $\varepsilon$ away in $\ell_1$…

Statistics Theory · Mathematics 2021-10-14 Sudeep Salgia , Qing Zhao , Lang Tong

Flexible variational distributions improve variational inference but are harder to optimize. In this work we present a control variate that is applicable for any reparameterizable distribution with known mean and covariance matrix, e.g.…

Machine Learning · Computer Science 2020-10-26 Tomas Geffner , Justin Domke

This paper proposes a new probabilistic classification algorithm using a Markov random field approach. The joint distribution of class labels is explicitly modelled using the distances between feature vectors. Intuitively, a class label…

Computation · Statistics 2010-06-02 Nial Friel , Anthony N. Pettitt

Sensitivity properties describe how changes to the input of a program affect the output, typically by upper bounding the distance between the outputs of two runs by a monotone function of the distance between the corresponding inputs. When…

Logic in Computer Science · Computer Science 2020-08-11 Alejandro Aguirre , Gilles Barthe , Justin Hsu , Benjamin Lucien Kaminski , Joost-Pieter Katoen , Christoph Matheja

The most widely used internal measure for clustering evaluation is the silhouette coefficient, whose naive computation requires a quadratic number of distance calculations, which is clearly unfeasible for massive datasets. Surprisingly,…

Data Structures and Algorithms · Computer Science 2021-01-21 Federico Altieri , Andrea Pietracaprina , Geppino Pucci , Fabio Vandin

Curve matching is a prediction technique that relies on predictive mean matching, which matches donors that are most similar to a target based on the predictive distance. Even though this approach leads to high prediction accuracy, the…

Methodology · Statistics 2022-07-12 Anaïs Fopma , Mingyang Cai , Stef van Buuren , Gerko Vink

It is well-known that the scale-free networks are ubiquitous in nature and society and have been one of the hotspot topic in complex networks. Recently, scholars presented a large quantity of scale-free networks by calculating cumulative…

Social and Information Networks · Computer Science 2020-11-02 Xiaomin Wang , Bing Yao

Diffusions are a successful technique to sample from high-dimensional distributions. The target distribution can be either explicitly given or learnt from a collection of samples. They implement a diffusion process whose endpoint is a…

Machine Learning · Computer Science 2025-09-03 Andrea Montanari

If the prior probability distributions of all possible hypothetical true means and all possible observed means of a continuous variable are conditional on the universal set of all numbers (i.e., before the nature of a study is known and a…

Methodology · Statistics 2025-06-05 Huw Llewelyn

Some properties of chaotic dynamical systems can be probed through features of recurrences, also called analogs. In practice, analogs are nearest neighbours of the state of a system, taken from a large database called the catalog. Analogs…

Dynamical Systems · Mathematics 2021-10-27 Paul Platzer , Pascal Yiou , Philippe Naveau , Jean-François Filipot , Maxime Thiebaut , Pierre Tandeo

Probabilistic programs are typically normal-looking programs describing posterior probability distributions. They intrinsically code up randomized algorithms and have long been at the heart of modern machine learning and approximate…

Programming Languages · Computer Science 2023-02-14 Lutz Klinkenberg , Tobias Winkler , Mingshuai Chen , Joost-Pieter Katoen

We analyze the distribution of the distance between two nodes, sampled uniformly at random, in digraphs generated via the directed configuration model, in the supercritical regime. Under the assumption that the covariance between the…

Probability · Mathematics 2017-04-24 Pim van der Hoorn , Mariana Olvera-Cravioto

Observed clusters should be modelled by considering the distribution function to be a random variable that quantifies the degree of excitation of the system's normal modes. A system of canonical coordinates for the space of DFs is…

Astrophysics of Galaxies · Physics 2021-08-11 Jun Yan Lau , James Binney

In this paper, we address the question of comparison between populations of trees. We study an statistical test based on the distance between empirical mean trees, as an analog of the two sample z statistic for comparing two means. Despite…

Statistics Theory · Mathematics 2007-08-14 Ana Georgina Flesia , Ricardo Fraiman

In this work, we give a novel general approach for distribution testing. We describe two techniques: our first technique gives sample-optimal testers, while our second technique gives matching sample lower bounds. As a consequence, we…

Data Structures and Algorithms · Computer Science 2016-05-10 Ilias Diakonikolas , Daniel M. Kane

A theoretical framework is developed to describe the transformation that distributes probability density functions uniformly over space. In one dimension, the cumulative distribution can be used, but does not generalize to higher…

Neural and Evolutionary Computing · Computer Science 2016-09-08 Eric Kee

In some fields of applications of stable distributions, especially in economics, it appears, that data have distributions similar to stable in a large region, but do not have such heavy tails. Our aim in this note is to propose several…

Probability · Mathematics 2014-03-17 Lenka Slámová , Lev B. Klebanov

Clustering is an important part of many modern data analysis pipelines, including network analysis and data retrieval. There are many different clustering algorithms developed by various communities, and it is often not clear which…

Machine Learning · Computer Science 2019-10-04 Maria-Florina Balcan , Travis Dick , Manuel Lang

We revisit extending the Kolmogorov-Smirnov distance between probability distributions to the multidimensional setting and make new arguments about the proper way to approach this generalization. Our proposed formulation maximizes the…

Computation · Statistics 2025-04-16 Peter Matthew Jacobs , Foad Namjoo , Jeff M. Phillips
‹ Prev 1 8 9 10 Next ›