相关论文: Magnitude Distance: A Geometric Measure of Dataset…
The ability to mimic human notions of semantic distance has widespread applications. Some measures rely only on raw text (distributional measures) and some rely on knowledge sources such as WordNet. Although extensive studies have been…
Common measures of neural representational (dis)similarity are designed to be insensitive to rotations and reflections of the neural activation space. Motivated by the premise that the tuning of individual units may be important, there has…
The concept of dimension is essential to grasp the complexity of data. A naive approach to determine the dimension of a dataset is based on the number of attributes. More sophisticated methods derive a notion of intrinsic dimension (ID)…
Clustering in high dimension spaces is a difficult task; the usual distance metrics may no longer be appropriate under the curse of dimensionality. Indeed, the choice of the metric is crucial, and it is highly dependent on the dataset…
Magnitude is a numerical invariant of metric spaces introduced by Leinster, motivated by considerations from category theory. This paper extends the original definition for finite spaces to compact spaces, in an equivalent but more natural…
Distance function is a main metrics of measuring the affinity between two data points in machine learning. Extant distance functions often provide unreachable distance values in real applications. This can lead to incorrect measure of the…
Machine learning has been proven to be effective in various application areas, such as object and speech recognition on mobile systems. Since a critical key to machine learning success is the availability of large training data, many…
Bias evaluation is fundamental to trustworthy AI, both in terms of checking data quality and in terms of checking the outputs of AI systems. In testing data quality, for example, one may study the distance of a given dataset, viewed as a…
Magnitude is a numerical invariant of metric spaces and graphs, analogous, in a precise sense, to Euler characteristic. Magnitude homology is an algebraic invariant constructed to categorify magnitude. Among the important features of the…
The study of the topological structure of complex networks has fascinated researchers for several decades, and today we have a fairly good understanding of the types and reoccurring characteristics of many different complex networks.…
The magnitude of metric spaces does not appear to possess a simple, convenient continuity property, and previous studies have presented affirmative results under additional constraints or weaker notions, as well as counterexamples. In this…
Distance-based classification is among the most competitive classification methods for time series data. The most critical component of distance-based classification is the selected distance function. Past research has proposed various…
The rapid growth of high-dimensional datasets across various scientific domains has created a pressing need for new statistical methods to compare distributions supported on their underlying structures. Assessing similarity between datasets…
Graph comparison plays a major role in many network applications. We often need a similarity metric for comparing networks according to their structural properties. Various network features - such as degree distribution and clustering…
Magnitude is a numerical invariant of enriched categories, including in particular metric spaces as $[0,\infty)$-enriched categories. We show that in many cases magnitude can be categorified to a homology theory for enriched categories,…
Magnitude homology is an $\mathbf{R}^+$-graded homology theory of metric spaces that captures information on the complexity of geodesics. Here we address the question: when are two metric spaces magnitude homology equivalent, in the sense…
In this paper we introduce the persistent magnitude, a new numerical invariant of (sufficiently nice) graded persistence modules. It is a weighted and signed count of the bars of the persistence module, in which a bar of the form $[a,b)$ in…
The most useful data mining primitives are distance measures. With an effective distance measure, it is possible to perform classification, clustering, anomaly detection, segmentation, etc. For single-event time series Euclidean Distance…
The ability to represent and compare machine learning models is crucial in order to quantify subtle model changes, evaluate generative models, and gather insights on neural network architectures. Existing techniques for comparing data…
Distance metric learning (DML) has been studied extensively in the past decades for its superior performance with distance-based algorithms. Most of the existing methods propose to learn a distance metric with pairwise or triplet…