Related papers: Classifying token frequencies using angular Minkow…
In a previous article (Rahman, Guns, Rousseau, and Engels, 2015) we described several approaches to determine the cognitive distance between two units. One of these approaches was based on what we called barycenters in N dimensions. The…
Distance correlation is a novel class of multivariate dependence measure, taking positive values between 0 and 1, and applicable to random vectors of arbitrary dimensions, not necessarily equal. It offers several advantages over the…
In deep regression, capturing the relationship among continuous labels in feature space is a fundamental challenge that has attracted increasing interest. Addressing this issue can prevent models from converging to suboptimal solutions…
We propose a new method for estimating the intrinsic dimension of a dataset by applying the principle of regularized maximum likelihood to the distances between close neighbors. We propose a regularization scheme which is motivated by…
It is a classical fact, that given an arbitrary n-dimensional convex body, there exists an appropriate sequence of Minkowski symmetrizations (or Steiner symmetrizations), that converges in Hausdorff metric to a Euclidean ball. Here we…
The local theory of complex dimensions describes the oscillations in the geometry (spectra and dynamics) of fractal strings. Such geometric oscillations can be seen most clearly in the explicit volume formula for the tubular neighborhoods…
This paper considers the problem of distinguishing between classical and quantum domains in macroscopic phenomena using tests based on probability and it presents a condition on the ratios of the outcomes being the same (Ps) to being…
Local learning of sparse image models has proven to be very effective to solve inverse problems in many computer vision applications. To learn such models, the data samples are often clustered using the K-means algorithm with the Euclidean…
Starting with a similarity function between objects, it is possible to define a distance metric on pairs of objects, and more generally on probability distributions over them. These distance metrics have a deep basis in functional analysis,…
We give a general unified method that can be used for $L_1$ {\em closeness testing} of a wide range of univariate structured distribution families. More specifically, we design a sample optimal and computationally efficient algorithm for…
In the context of clustering, we consider a generative model in a Euclidean ambient space with clusters of different shapes, dimensions, sizes and densities. In an asymptotic setting where the number of points becomes large, we obtain…
Minkowski Tensors are tensorial generalizations of the scalar Minkowski Functionals. Due to their tensorial nature they contain additional morphological information of structures, in particular about shape and alignment, in comparison to…
In recent years, the learned local descriptors have outperformed handcrafted ones by a large margin, due to the powerful deep convolutional neural network architectures such as L2-Net [1] and triplet based metric learning [2]. However,…
We introduce a new measure of distance between datasets, based on vineyards from topological data analysis, which we call the vineyard distance. Vineyard distance measures the extent of topological change along an interpolation from one…
Gravitational waves from inspiraling compact objects provide us with information of the distance scale since we can infer the absolute luminosity of the source from analysis of the wave form, which is known as standard sirens. The first…
A new clustering accuracy measure is proposed to determine the unknown number of clusters and to assess the quality of clustering of a data set given in any dimensional space. Our validity index applies the classical nonparametric…
The most useful data mining primitives are distance measures. With an effective distance measure, it is possible to perform classification, clustering, anomaly detection, segmentation, etc. For single-event time series Euclidean Distance…
Comparing metric measure spaces (i.e. a metric space endowed with aprobability distribution) is at the heart of many machine learning problems. The most popular distance between such metric measure spaces is theGromov-Wasserstein (GW)…
Time series clustering is an unsupervised learning method for classifying time series data into groups with similar behavior. It is used in applications such as healthcare, finance, economics, energy, and climate science. Several time…
Prior methods for retrieval of nearest neighbors in high dimensions are fast and approximate--providing probabilistic guarantees of returning the correct answer--or slow and exact performing an exhaustive search. We present Certified…