English
Related papers

Related papers: Characterizing how 'distributional' NLP corpora di…

200 papers

We provide an implementation to compute the flat metric in any dimension. The flat metric, also called dual bounded Lipschitz distance, generalizes the well-known Wasserstein distance $W_1$ to the case that the distributions are of unequal…

Machine Learning · Computer Science 2025-06-17 Henri Schmidt , Christian Düll

We propose a fundamental metric for measuring the distance between two distributions. This metric, referred to as the decision-focused (DF) divergence, is tailored to stochastic linear optimization problems in which the objective…

Statistics Theory · Mathematics 2026-02-04 Suhan Liu , Mo Liu

This article provides an overview on the statistical modeling of complex data as increasingly encountered in modern data analysis. It is argued that such data can often be described as elements of a metric space that satisfies certain…

Methodology · Statistics 2024-02-28 Paromita Dubey , Yaqing Chen , Hans-Georg Müller

Generative models are invaluable in many fields of science because of their ability to capture high-dimensional and complicated distributions, such as photo-realistic images, protein structures, and connectomes. How do we evaluate the…

Understanding geometric properties of natural language processing models' latent spaces allows the manipulation of these properties for improved performance on downstream tasks. One such property is the amount of data spread in a model's…

Machine Learning · Computer Science 2023-08-02 Anna C. Marbut , Katy McKinney-Bock , Travis J. Wheeler

With the increasing availability of data objects in the form of probability distributions, there is a growing need for statistical methods tailored to distributional data. Distance measures, especially the pairwise distance matrix between…

Methodology · Statistics 2026-03-03 Edward Shao , Junyoung Park , Naresh Punjabi , Hui Jiang , Irina Gaynanova

This paper presents methods to compare high order networks, defined as weighted complete hypergraphs collecting relationship functions between elements of tuples. They can be considered as generalizations of conventional networks where only…

Social and Information Networks · Computer Science 2016-01-20 Weiyu Huang , Alejandro Ribeiro

It is well-understood that different algorithms, training processes, and corpora produce different word embeddings. However, less is known about the relation between different embedding spaces, i.e. how far different sets of embeddings…

Computation and Language · Computer Science 2020-05-19 Xuhui Zhou , Zaixiang Zheng , Shujian Huang

This article inspects whether a multivariate distribution is different from a specified distribution or not, and it also tests the equality of two multivariate distributions. In the course of this study, a graphical tool-kit using…

Methodology · Statistics 2024-08-19 Pratim Guha Niyogi , Subhra Sankar Dhar

In machine learning, observation features are measured in a metric space to obtain their distance function for optimization. Given similar features that are statistically sufficient as a population, a statistical distance between two…

Machine Learning · Statistics 2020-06-23 Xin Lu

In a recent paper, Aldous, Blanc and Curien asked which distributions can be expressed as the distance between two independent random variables on some separable measured metric space. We show that every nonnegative discrete distribution…

Probability · Mathematics 2025-07-22 Renan Gross

Common measures of neural representational (dis)similarity are designed to be insensitive to rotations and reflections of the neural activation space. Motivated by the premise that the tuning of individual units may be important, there has…

Machine Learning · Computer Science 2023-11-17 Meenakshi Khosla , Alex H. Williams

Distance distributions are a key building block in stochastic geometry modelling of wireless networks and in many other fields in mathematics and science. In this paper, we propose a novel framework for analytically computing the closed…

Information Theory · Computer Science 2019-03-20 Ross Pure , Salman Durrani , Fei Tong , Jianping Pan

We study the Hausdorff distance between a random polytope, defined as the convex hull of i.i.d. random points, and the convex hull of the support of their distribution. As particular examples, we consider uniform distributions on convex…

Statistics Theory · Mathematics 2018-07-05 Victor-Emmanuel Brunel

Testing the equality of two conditional distributions is crucial in various modern applications, including transfer learning and causal inference. Despite its importance, this fundamental problem has received surprisingly little attention…

Methodology · Statistics 2025-09-04 Jian Yan , Zhuoxi Li , Xianyang Zhang

Graphs are used in almost every scientific discipline to express relations among a set of objects. Algorithms that compare graphs, and output a closeness score, or a correspondence among their nodes, are thus extremely important. Despite…

Discrete Mathematics · Computer Science 2020-11-17 Sam Safavi , José Bento

The empirical distribution function assigns mass $1/n$ to each of the $n$ observations in a sample. As these are highly variable, estimation error may be reduced by replacing them with estimated observations that are asymptotically less…

Methodology · Statistics 2026-05-26 Tommaso Lando , Lorenzo Tedesco

We examine the geometry of the spaces between particles in diffusion-limited cluster aggregation, a numerical model of aggregating suspensions. Computing the distribution of distances from each point to the nearest particle, we show that it…

Statistical Mechanics · Physics 2009-11-07 R. M. L. Evans , M. D. Haw

In this paper a new dissimilarity measure to identify groups of assets dynamics is proposed. The underlying generating process is assumed to be a diffusion process solution of stochastic differential equations and observed at discrete time.…

Statistical Finance · Quantitative Finance 2008-12-02 Alessandro De Gregorio , Stefano Maria Iacus

In this paper, we advocate a novel measure for the purpose of checking the quality of a cluster partition for a sample into several distinct classes, and thus, determine the unknown value for the true number of clusters prevailing the…

Applications · Statistics 2024-04-12 Soumita Modak