English
Related papers

Related papers: Classifying token frequencies using angular Minkow…

200 papers

The cosmological luminosity-distance can be measured from gravitational wave (GW) standard sirens, free of astronomical distance ladders and the associated systematics. However, it may still contain systematics arising from various…

Cosmology and Nongalactic Astrophysics · Physics 2021-08-25 Pengjie Zhang , Hai Yu

K-Means clustering algorithm is one of the most commonly used clustering algorithms because of its simplicity and efficiency. K-Means clustering algorithm based on Euclidean distance only pays attention to the linear distance between…

Machine Learning · Computer Science 2022-06-13 Yiqun Zhang , Houbiao Li

Many methods in differentially private model training rely on computing the similarity between a query point (such as public or synthetic data) and private data. We abstract out this common subroutine and study the following fundamental…

Cryptography and Security · Computer Science 2024-03-15 Arturs Backurs , Zinan Lin , Sepideh Mahabadi , Sandeep Silwal , Jakub Tarnawski

For each given $p\in[1,\infty]$ we investigate certain sub-family $\mathcal{M}_p$ of the collection of all compact metric spaces $\mathcal{M}$ which are characterized by the satisfaction of a strengthened form of the triangle inequality…

Metric Geometry · Mathematics 2021-11-24 Facundo Mémoli , Zhengchao Wan

The Pearson distance has been advocated for improving the error performance of noisy channels with unknown gain and offset. The Pearson distance can only fruitfully be used for sets of $q$-ary codewords, called Pearson codes, that satisfy…

Information Theory · Computer Science 2016-11-17 Jos H. Weber , Kees A. Schouhamer Immink , Simon R. Blackburn

We introduce a variant of the $k$-nearest neighbor classifier in which $k$ is chosen adaptively for each query, rather than supplied as a parameter. The choice of $k$ depends on properties of each neighborhood, and therefore may…

Machine Learning · Computer Science 2019-05-31 Akshay Balsubramani , Sanjoy Dasgupta , Yoav Freund , Shay Moran

We analytically study proximity and distance properties of various kernels and similarity measures on graphs. This helps to understand the mathematical nature of such measures and can potentially be useful for recommending the adoption of…

Combinatorics · Mathematics 2018-08-17 Konstantin Avrachenkov , Pavel Chebotarev , Dmytro Rubanov

The paper introduces a new kernel-based Maximum Mean Discrepancy (MMD) statistic for measuring the distance between two distributions given finitely-many multivariate samples. When the distributions are locally low-dimensional, the proposed…

Machine Learning · Statistics 2018-09-03 Xiuyuan Cheng , Alexander Cloninger , Ronald R. Coifman

This paper presents a probabilistic generalization of the Generalized Optimal Sub-Pattern Assignment (GOSPA) metric, termed P-GOSPA. The GOSPA metric has been widely used to evaluate the distance between finite sets, particularly in…

Signal Processing · Electrical Eng. & Systems 2025-06-17 Yuxuan Xia , Ángel F. García-Fernández , Johan Karlsson , Kuo-Chu Chang , Ting Yuan , Lennart Svensson

Popular clustering algorithms based on usual distance functions (e.g., Euclidean distance) often suffer in high dimension, low sample size (HDLSS) situations, where concentration of pairwise distances has adverse effects on their…

Methodology · Statistics 2019-05-03 Soham Sarkar , Anil K. Ghosh

We have created a large database of similarity information between sub-regions of Hubble Space Telescope images. These data can be used to assess the accuracy of image search algorithms based on computer vision methods. The images were…

Instrumentation and Methods for Astrophysics · Physics 2025-04-25 Richard L. White , J. E. G. Peek

In contrast to a maximum-likelihood decoder, it is often desirable to use an incomplete decoder that can detect its decoding errors with high probability. One common choice is the bounded distance decoder. Bounds are derived for the total…

Information Theory · Computer Science 2012-07-26 Kenneth Andrews , Sam Dolinar

Nearest neighbor is a popular class of classification methods with many desirable properties. For a large data set which cannot be loaded into the memory of a single machine due to computation, communication, privacy, or ownership…

Machine Learning · Statistics 2019-11-01 Xingye Qiao , Jiexin Duan , Guang Cheng

The statistical distribution of galaxies is a powerful probe to constrain cosmological models and gravity. In particular the matter power spectrum $P(k)$ brings information about the cosmological distance evolution and the galaxy clustering…

Cosmology and Nongalactic Astrophysics · Physics 2017-06-20 J. -E. Campagne , J. Neveu , S. Plaszczynski

We provide a coherence-based approach to nonclassical behavior by means of distance measures. We develop a quantitative relation between coherence and nonclassicality quantifiers, which establish the nonclassicality as the maximum…

Quantum Physics · Physics 2022-07-20 Laura Ares , Alfredo Luis

In the Euclidean TSP with neighborhoods (TSPN), we are given a collection of n regions (neighborhoods) and we seek a shortest tour that visits each region. As a generalization of the classical Euclidean TSP, TSPN is also NP-hard. In this…

Computational Geometry · Computer Science 2017-03-07 Adrian Dumitrescu , Joseph S. B. Mitchell

To improve our understanding of connected systems, different tools derived from statistics, signal processing, information theory and statistical physics have been developed in the last decade. Here, we will focus on the graph comparison…

Physics and Society · Physics 2018-04-23 Johann H. Martínez , Mario Chavez

This paper defines a distance function that measures the dissimilarity between planar geometric figures formed with straight lines. This function can in turn be used in partial matching of different geometric figures. For a given pair of…

Computer Vision and Pattern Recognition · Computer Science 2016-12-06 Apoorva Honnegowda Roopa , Shrisha Rao

Distance measures are part and parcel of many computer vision algorithms. The underlying assumption in all existing distance measures is that feature elements are independent and identically distributed. However, in real-world settings,…

Computer Vision and Pattern Recognition · Computer Science 2016-11-01 Muthukaruppan Swaminathan , Pankaj Kumar Yadav , Obdulio Piloto , Tobias Sjöblom , Ian Cheong

The goal of this thesis is to study the use of the Kantorovich-Rubinstein distance as to build a descriptor of sample complexity in classification problems. The idea is to use the fact that the Kantorovich-Rubinstein distance is a metric in…

Probability · Mathematics 2023-09-19 Gaël Giordano