Related papers: Measuring spatial uniformity with the hypersphere …
Unimodality constitutes a key property indicating grouping behavior of the data around a single mode of its density. We propose a method that partitions univariate data into unimodal subsets through recursive splitting around valley points…
In this paper, the performance of non-uniform spacing of sensors is evaluated for the MUSIC algorithm which estimates the direction of arrival (DOA) of a narrowband plane wave impinging on an array of sensors. Unlike uniform sensor spacing…
Phylogenetic inference-the derivation of a hypothesis for the common evolutionary history of a group of species- is an active area of research at the intersection of biology, computer science, mathematics, and statistics. One assumes the…
Accurately estimating data density is crucial for making informed decisions and modeling in various fields. This paper presents a novel nonparametric density estimation procedure that utilizes bivariate penalized spline smoothing over…
A novel method to obtain hierarchical and overlapping clusters from network data -i.e., a set of nodes endowed with pairwise dissimilarities- is presented. The introduced method is hierarchical in the sense that it outputs a nested…
The Ultraviolet Near Infrared Optical Northern Survey (UNIONS) is a photometric survey in the northern sky. The quality of the data in the $r$ band provides precise shape measurements to measure the growth of structures using cosmic shear.…
Clustering is spotting pattern in a group of objects and resultantly grouping the similar objects together. Objects have attributes which are not always numerical, sometimes attributes have domain or categories to which they could belong…
Persistent homology, the study of holes that appear in data as one thickens balls centered around its points over time, has theoretically guaranteed stability. That is, small data perturbations guarantee small changes in the lifetimes of…
In absence of a lens to form an image, incoherent or partially coherent light scattering off an obstructive or reflective object forms a broad intensity distribution in the far field with only feeble spatial features. We show here that…
Unsupervised machine learning, and in particular data clustering, is a powerful approach for the analysis of datasets and identification of characteristic features occurring throughout a dataset. It is gaining popularity across scientific…
Existing online change-point detection (CPD) methods rely on fixed-dimensional Euclidean summaries, implicitly assuming that distributional changes are well captured by moment-based or feature-based representations. They can obscure…
Hyperuniform particle arrangements are characterized by a local number variance that grows more slowly than the volume of the observation window. We generalize this concept to describe particle systems in which particles carry weights:…
We survey the emerging area of compression-based, parameter-free, similarity distance measures useful in data-mining, pattern recognition, learning and automatic semantics extraction. Given a family of distances on a set of objects, a…
Melodic similarity measurement is of key importance in music information retrieval. In this paper, we use geometric matching techniques to measure the similarity between two melodies. We represent music as sets of points or sets of…
This paper relaxes the restrictive symmetry conditions adopted in [4], [5] and extends their universal feature selection framework to accommodate noisy observations as well as attribute structures that may exhibit directional preferences.…
Data-dependent metrics are powerful tools for learning the underlying structure of high-dimensional data. This article develops and analyzes a data-dependent metric known as diffusion state distance (DSD), which compares points using a…
The local number variance associated with a spherical sampling window of radius $R$ enables a classification of many-particle systems in $d$-dimensional Euclidean space according to the degree to which large-scale density fluctuations are…
Rotationally symmetric distributions on the p-dimensional unit hypersphere, extremely popular in directional statistics, involve a location parameter theta that indicates the direction of the symmetry axis. The most classical way of…
This paper provides a new similarity detection algorithm. Given an input set of multi-dimensional data points, where each data point is assumed to be multi-dimensional, and an additional reference data point for similarity finding, the…
Maximum Mean Discrepancy (MMD) has been widely used in the areas of machine learning and statistics to quantify the distance between two distributions in the $p$-dimensional Euclidean space. The asymptotic property of the sample MMD has…