English
Related papers

Related papers: Characterizing how 'distributional' NLP corpora di…

200 papers

In distributed and federated learning, heterogeneity across data sources remains a major obstacle to effective model aggregation and convergence. We focus on feature heterogeneity and introduce energy distance as a sensitive measure for…

Machine Learning · Statistics 2025-01-28 Mengchen Fan , Baocheng Geng , Roman Shterenberg , Joseph A. Casey , Zhong Chen , Keren Li

In this paper we propose and study a class of nonparametric, yet interpretable measures of association between two random vectors $X$ and $Y$ taking values in $\mathbb{R}^{d_1}$ and $\mathbb{R}^{d_2}$ respectively ($d_1, d_2\ge 1$). These…

Statistics Theory · Mathematics 2024-11-21 Nabarun Deb , Promit Ghosal , Bodhisattva Sen

We propose a new method for local distance metric learning based on sample similarity as side information. These local metrics, which utilize conical combinations of metric weight matrices, are learned from the pooled spatial…

Machine Learning · Computer Science 2019-02-25 YInjie Huang , Cong Li , Michael Georgiopoulos , Georgios C. Anagnostopoulos

We study a general framework of distributional computational graphs: computational graphs whose inputs are probability distributions rather than point values. We analyze the discretization error that arises when these graphs are evaluated…

Machine Learning · Statistics 2026-02-13 Olof Hallqvist Elias , Michael Selby , Phillip Stanley-Marbell

We introduce a novel, geometry-aware distance metric for the family of von Mises-Fisher (vMF) distributions, which are fundamental models for directional data on the unit hypersphere. Although the vMF distribution is widely employed in a…

Machine Learning · Statistics 2025-04-22 Kisung You , Dennis Shung , Mauro Giuffrè

Given $M \geq 2$ distributions defined on a general measurable space, we introduce a nonparametric (kernel) measure of multi-sample dissimilarity (KMD) -- a parameter that quantifies the difference between the $M$ distributions. The…

Statistics Theory · Mathematics 2022-10-18 Zhen Huang , Bodhisattva Sen

This paper concerns a method of testing equality of distribution of random convex compact sets and the way how to use the test to distinguish between two realisations of general random sets. The family of metrics on the space of…

Statistics Theory · Mathematics 2018-01-09 Vesna Gotovac , Kateřina Helisová

We study fractal measures on Euclidean space through the dynamics of "zooming in" on typical points. The resulting family of measures (the "scenery"), can be interpreted as an orbit in an appropriate dynamical system which often…

Dynamical Systems · Mathematics 2013-07-31 Michael Hochman

The ability to represent and compare machine learning models is crucial in order to quantify subtle model changes, evaluate generative models, and gather insights on neural network architectures. Existing techniques for comparing data…

In this work we study systems consisting of a group of moving particles. In such systems, often some important parameters are unknown and have to be estimated from observed data. Such parameter estimation problems can often be solved via a…

Applications · Statistics 2023-07-11 Chen Cheng , Linjie Wen , Jinglai Li

The assessment of segmentation quality plays a fundamental role in the development, optimization, and comparison of segmentation methods which are used in a wide range of applications. With few exceptions, quality assessment is performed…

Computer Vision and Pattern Recognition · Computer Science 2025-04-02 Niklas Rottmayer , Claudia Redenbach

In this work, we empirically explore the question: how can we assess the quality of samples from some target distribution? We assume that the samples are provided by some valid Monte Carlo procedure, so we are guaranteed that the collection…

Machine Learning · Computer Science 2016-06-21 Arjumand Masood , Weiwei Pan , Finale Doshi-Velez

The notion of concept drift refers to the phenomenon that the distribution, which is underlying the observed data, changes over time; as a consequence machine learning models may become inaccurate and need adjustment. Many unsupervised…

Machine Learning · Computer Science 2022-02-22 Fabian Hinder , Valerie Vaquet , Barbara Hammer

In wireless networks, the knowledge of nodal distances is essential for several areas such as system configuration, performance analysis and protocol design. In order to evaluate distance distributions in random networks, the underlying…

Information Theory · Computer Science 2012-01-24 Sunil Srinivasa , Martin Haenggi

The aim of the present article is to introduce a concept which allows to generalise the notion of Poissonian pair correlation, a second-order equidistribution property, to higher dimensions. Roughly speaking, in the one-dimensional setting,…

Number Theory · Mathematics 2018-09-18 Aicke Hinrichs , Lisa Kaltenböck , Gerhard Larcher , Wolfgang Stockinger , Mario Ullrich

Distance queries are a basic tool in data analysis. They are used for detection and localization of change for the purpose of anomaly detection, monitoring, or planning. Distance queries are particularly useful when data sets such as…

Data Structures and Algorithms · Computer Science 2015-03-20 Edith Cohen

Heterogeneity in dynamics in the form of non-Gaussian molecular displacement distributions appears ubiquitously in soft matter. We address the quantification of such heterogeneity using an information-theoretic measure of the distance…

Soft Condensed Matter · Physics 2020-08-04 Rahul Dandekar , Soumyakanti Bose , Suman Dutta

Distribution testing deals with what information can be deduced about an unknown distribution over $\{1,\ldots,n\}$, where the algorithm is only allowed to obtain a relatively small number of independent samples from the distribution. In…

Computational Complexity · Computer Science 2016-09-23 Eldar Fischer , Oded Lachish , Yadu Vasudev

The purpose of unconditional text generation is to train a model with real sentences, then generate novel sentences of the same quality and diversity as the training data. However, when different metrics are used for comparing the methods…

Computation and Language · Computer Science 2020-07-03 Ping Cai , Xingyuan Chen , Peng Jin , Hongjun Wang , Tianrui Li

We develop a kernel projected Wasserstein distance for the two-sample test, an essential building block in statistics and machine learning: given two sets of samples, to determine whether they are from the same distribution. This method…

Statistics Theory · Mathematics 2022-05-10 Jie Wang , Rui Gao , Yao Xie