English
Related papers

Related papers: Classifying token frequencies using angular Minkow…

200 papers

Metric based comparison operations such as finding maximum, nearest and farthest neighbor are fundamental to studying various clustering techniques such as $k$-center clustering and agglomerative hierarchical clustering. These techniques…

Data Structures and Algorithms · Computer Science 2021-05-13 Raghavendra Addanki , Sainyam Galhotra , Barna Saha

Pairwise Euclidean distance calculation is a fundamental step in many machine learning and data analysis algorithms. In real-world applications, however, these distances are frequently distorted by heteroskedastic noise$\unicode{x2014}$a…

Machine Learning · Statistics 2025-09-12 Keyi Li , Yuval Kluger , Boris Landa

The use of correntropy as a similarity measure has been increasing in different scenarios due to the well-known ability to extract high-order statistic information from data. Recently, a new similarity measure between complex random…

Information Theory · Computer Science 2017-10-03 João Guimarães

The conditional mutual information quantifies the conditional dependence of two random variables. It has numerous applications; it forms, for example, part of the definition of transfer entropy, a common measure of the causal relationship…

Information Theory · Computer Science 2024-04-15 Jake Witter , Conor Houghton

Optimal transport and its related problems, including optimal partial transport, have proven to be valuable tools in machine learning for computing meaningful distances between probability or positive measures. This success has led to a…

Machine Learning · Computer Science 2023-07-26 Xinran Liu , Yikun Bai , Huy Tran , Zhanqi Zhu , Matthew Thorpe , Soheil Kolouri

Accurate approximation of probability measures is essential in numerical applications. This paper explores the quantization of probability measures using the maximum mean discrepancy (MMD) distance as a guiding metric. We first investigate…

Optimization and Control · Mathematics 2025-03-18 Zahra Mehraban , Alois Pichler

When data is of an extraordinarily large size or physically stored in different locations, the distributed nearest neighbor (NN) classifier is an attractive tool for classification. We propose a novel distributed adaptive NN classifier for…

Machine Learning · Statistics 2023-06-06 Ruiqi Liu , Ganggang Xu , Zuofeng Shang

Approximate Bayesian computation performs approximate inference for models where likelihood computations are expensive or impossible. Instead simulations from the model are performed for various parameter values and accepted if they are…

Computation · Statistics 2015-12-16 Dennis Prangle

Optical atomic clocks are our most precise tools to measure time and frequency. They enable precision frequency comparisons between atoms in separate locations to probe the space-time variation of fundamental constants, the properties of…

Atomic Physics · Physics 2022-09-09 B. C. Nichol , R. Srinivas , D. P. Nadlinger , P. Drmota , D. Main , G. Araneda , C. J. Ballance , D. M. Lucas

A suitable scalar metric can help measure multi-calibration, defined as follows. When the expected values of observed responses are equal to corresponding predicted probabilities, the probabilistic predictions are known as "perfectly…

Methodology · Statistics 2026-04-17 Ido Guy , Daniel Haimovich , Fridolin Linder , Nastaran Okati , Lorenzo Perini , Niek Tax , Mark Tygert

We examine the class of weakly porous sets in Euclidean spaces. As our first main result we show that the distance weight $w(x)=\operatorname{dist}(x,E)^{-\alpha}$ belongs to the Muckenhoupt class $A_1$, for some $\alpha>0$, if and only if…

Classical Analysis and ODEs · Mathematics 2024-07-19 Theresa C. Anderson , Juha Lehrbäck , Carlos Mudarra , Antti V. Vähäkangas

Collections of probability distributions arise in a variety of applications ranging from user activity pattern analysis to brain connectomics. In practice these distributions can be defined over diverse domain types including finite…

Methodology · Statistics 2023-06-16 Raif Rustamov , Subhabrata Majumdar

The Minkowski content of a compact set is a fine measure of its geometric scaling. For Lebesgue null sets it measures the decay of the Lebesgue measure of epsilon neighbourhoods of the set. It is well known that self-similar sets,…

Dynamical Systems · Mathematics 2023-03-14 Sascha Troscheit

The main contribution of this dissertation is the introduction of new or improved approximation algorithms and data structures for several similarity search problems. We examine the furthest neighbor query, the annulus query, distance…

Data Structures and Algorithms · Computer Science 2019-06-13 Johan von Tangen Sivertsen

Information distance is a parameter-free similarity measure based on compression, used in pattern recognition, data mining, phylogeny, clustering, and classification. The notion of information distance is extended from pairs to multiples…

Computer Vision and Pattern Recognition · Computer Science 2009-05-21 Paul M. B. Vitanyi

Although recovering an Euclidean distance matrix from noisy observations is a common problem in practice, how well this could be done remains largely unknown. To fill in this void, we study a simple distance matrix estimate based upon the…

Machine Learning · Statistics 2014-09-18 Luwan Zhang , Grace Wahba , Ming Yuan

Existing sequence alignment algorithms use heuristic scoring schemes which cannot be used as objective distance metrics. Therefore one relies on measures like the p- or log-det distances, or makes explicit, and often simplistic, assumptions…

Genomics · Quantitative Biology 2015-05-19 Orion Penner , Peter Grassberger , Maya Paczuski

Given a probability measure with density, Fermat distances and density-driven metrics are conformal transformations of the Euclidean metric that shrink distances in high density areas and enlarge distances in low density areas. Although…

Statistics Theory · Mathematics 2026-01-22 Jérôme Taupin , Frédéric Chazal

The K-Nearest Neighbors (KNN) algorithm is widely used for classification and regression; however, it suffers from limitations, including the equal treatment of all samples. We propose Information-Modified KNN (IM-KNN), a novel approach…

Machine Learning · Computer Science 2025-07-11 Mohammad Ali Vahedifar , Azim Akhtarshenas , Mohammad Mohammadi Rafatpanah , Maryam Sabbaghian

We investigate the statistical task of closeness (or equivalence) testing for multidimensional distributions. Specifically, given sample access to two unknown distributions $\mathbf p, \mathbf q$ on $\mathbb R^d$, we want to distinguish…

Data Structures and Algorithms · Computer Science 2023-11-23 Ilias Diakonikolas , Daniel M. Kane , Sihan Liu
‹ Prev 1 8 9 10 Next ›