English
Related papers

Related papers: Characterizing how 'distributional' NLP corpora di…

200 papers

Computing the similarity between two probability distributions is a recurring theme across control. We introduce a unified family of distances between the probability distributions of two random variables that is based on the discrepancy…

Systems and Control · Electrical Eng. & Systems 2025-10-03 Alexandros E. Tzikas , Arec Jamgochian , Nazim Kemal Ure , Mykel J. Kochenderfer , Stephen P. Boyd

In this paper, we develop a finite mixture of convolutional distributions, a statistical model to analyze continuous data distributed approximately on a mixture of low-dimensional affine subspaces. The observations are assumed independent…

Statistics Theory · Mathematics 2026-04-21 Sunrit Chakraborty , XuanLong Nguyen

This work considers the asymptotic behavior of the distance between two sample covariance matrices (SCM). A general result is provided for a class of functionals that can be expressed as sums of traces of functions that are separately…

Statistics Theory · Mathematics 2023-12-25 Roberto Pereira , Xavier Mestre , David Gregoratti

The degree distribution is an important characteristic of complex networks. In many applications, quantification of degree distribution in the form of a fixed-length feature vector is a necessary step. On the other hand, we often need to…

Social and Information Networks · Computer Science 2013-12-24 Sadegh Aliakbary , Jafar Habibi , Ali Movaghar

We develop techniques to quantify the degree to which a given (training or testing) example is an outlier in the underlying distribution. We evaluate five methods to score examples in a dataset by how well-represented the examples are, for…

Machine Learning · Computer Science 2019-10-30 Nicholas Carlini , Úlfar Erlingsson , Nicolas Papernot

We consider the problem of allocating samples to a finite set of discrete distributions in order to learn them uniformly well in terms of four common distance measures: $\ell_2^2$, $\ell_1$, $f$-divergence, and separation distance. To…

Machine Learning · Statistics 2019-12-10 Shubhanshu Shekhar , Tara Javidi , Mohammad Ghavamzadeh

Elliptically symmetric distributions are a classic example of a semiparametric model where the location vector and the scatter matrix (or a parameterization of them) are the two finite-dimensional parameters of interest, while the density…

Statistics Theory · Mathematics 2026-03-18 Stefano Fortunati , Jean-Pierre Delmas , Esa Ollila

The concentration of a distribution toward a lower bound is a conceptually simple property that closely relates to concepts of rarity and poverty, but that lacks a global descriptive statistic. We term this property 'shift' and define it as…

Methodology · Statistics 2025-06-16 Kenneth J. Locey , Brian D. Stein

Machine learning (ML) has employed various discretization methods to partition numerical attributes into intervals. However, an effective discretization technique remains elusive in many ML applications, such as association rule mining.…

Machine Learning · Computer Science 2023-11-07 Minakshi Kaushik , Rahul Sharma , Dirk Draheim

Distributed algorithms, particularly Diffusion Least Mean Square, are widely favored for their reliability, robustness, and fast convergence in various industries. However, limited observability of the target can compromise the integrity of…

Signal Processing · Electrical Eng. & Systems 2023-10-18 Mahdi Shamsi , Farokh Marvasti

Domain specific (dis-)similarity or proximity measures used e.g. in alignment algorithms of sequence data, are popular to analyze complex data objects and to cover domain specific data properties. Without an underlying vector space these…

Data Structures and Algorithms · Computer Science 2014-11-07 Andrej Gisbrecht , Frank-Michael Schleif

Which neural networks are similar is a fundamental question for both machine learning and neuroscience. Here, it is proposed to base comparisons on the predictive distributions of linear readouts from intermediate representations. In…

Machine Learning · Computer Science 2025-05-27 Heiko H. Schütt

We study the question of how reliably one can distinguish two quantum field theories (QFTs). Each QFT defines a probability distribution on the space of fields. The relative entropy provides a notion of proximity between these distributions…

High Energy Physics - Theory · Physics 2015-05-08 Vijay Balasubramanian , Jonathan J. Heckman , Alexander Maloney

To what extent can we distinguish one probability distribution from another? Are there quantitative measures of distinguishability? The goal of this tutorial is to approach such questions by introducing the notion of the "distance" between…

Data Analysis, Statistics and Probability · Physics 2015-06-23 Ariel Caticha

When we represent a network of sensors in Euclidean space by a graph, there are two distances between any two nodes that we may consider. One of them is the Euclidean distance. The other is the distance between the two nodes in the graph,…

Networking and Internet Architecture · Computer Science 2009-06-10 Rodrigo S. C. Leao , Valmir C. Barbosa

There are many open questions pertaining to the statistical analysis of random objects, which are increasingly encountered. A major challenge is the absence of linear operations in such spaces. A basic statistical task is to quantify…

Methodology · Statistics 2025-08-01 Wookyeong Song , Hans-Georg Müller

Distance metric learning is a successful way to enhance the performance of the nearest neighbor classifier. In most cases, however, the distribution of data does not obey a regular form and may change in different parts of the feature…

Computer Vision and Pattern Recognition · Computer Science 2018-03-19 Hossein Rajabzadeh , Mansoor Zolghadri Jahromi , Mohammad Sadegh Zare , Mostafa Fakhrahmad

Network data arises through observation of relational information between a collection of entities. Recent work in the literature has independently considered when (i) one observes a sample of networks, connectome data in neuroscience being…

Methodology · Statistics 2022-06-22 George Bolt , Simón Lunagómez , Christopher Nemeth

The ability to measure similarity between documents enables intelligent summarization and analysis of large corpora. Past distances between documents suffer from either an inability to incorporate semantic similarities between words or from…

Machine Learning · Computer Science 2019-11-05 Mikhail Yurochkin , Sebastian Claici , Edward Chien , Farzaneh Mirzazadeh , Justin Solomon

In this paper we introduce the idea of partially sorting data to design nonparametric tests. This approach gives rise to tests that are sensitive to both the order and the underlying distribution of the data. We focus in particular on a…

Statistics Theory · Mathematics 2022-10-27 Krzysztof Bisewski , H. M. Jansen , Yoni Nazarathy