Related papers: Field Formulation of Parzen Data Analysis
The density of electron-hole pairs produced in a graphene sample immersed in a homogeneous time-dependent electrical field is evaluated. Because low energy charge carriers in graphene are described by relativistic quantum mechanics, the…
Clustering in image analysis is a central technique that allows to classify elements of an image. We describe a simple clustering technique that uses the method of similarity matrices. We expand upon recent results in spectral analysis for…
The observed abundance of high-redshift galaxies and clusters contains precious information about the properties of the initial perturbations. We present a method to compute analytically the number density of objects as a function of mass…
This paper explores the critical role of data clustering in data science, emphasizing its methodologies, tools, and diverse applications. Traditional techniques, such as partitional and hierarchical clustering, are analyzed alongside…
A fundamental drawback of kernel-based statistical models is their limited scalability to large data sets, which requires resorting to approximations. In this work, we focus on the popular Gaussian kernel and on techniques to linearize…
DBSCAN is a classical density-based clustering procedure with tremendous practical relevance. However, DBSCAN implicitly needs to compute the empirical density for each sample point, leading to a quadratic worst-case time complexity, which…
Kernel density estimation is a popular method for estimating unseen probability distributions. However, the convergence of these classical estimators to the true density slows down in high dimensions. Moreover, they do not define meaningful…
Measurements of the cluster abundance as a function of mass and redshift provide an important cosmological test that probe not only the expansion rate but also the growth of perturbations. In this paper we adopt a scalar field scenario…
We propose a field-theoretical approach to a polymer system immersed in an ideal mixture of clustering centers. The system contains several species of these clustering centers with different functionality, each of which connects a fixed…
We formulate an optimization problem to estimate probability densities in the context of multidimensional problems that are sampled with uneven probability. It considers detector sensitivity as an heterogeneous density and takes advantage…
The simultaneous grouping of rows and columns is an important technique that is increasingly used in large-scale data analysis. In this paper, we present a novel co-clustering method using co-variables in its construction. It is based on a…
Extracting an understanding of the underlying system from high dimensional data is a growing problem in science. Discovering informative and meaningful features is crucial for clustering, classification, and low dimensional data embedding.…
A random Gaussian density field contains a fixed amount of Fisher information on the amplitude of its power spectrum. For a given smoothing scale, however, that information is not evenly distributed throughout the smoothed field. We…
We propose a novel approach to the problem of multilevel clustering, which aims to simultaneously partition data in each group and discover grouping patterns among groups in a potentially large hierarchically structured corpus of data. Our…
Discrete mixture models are one of the most successful approaches for density estimation. Under a Bayesian nonparametric framework, Dirichlet process location-scale mixture of Gaussian kernels is the golden standard, both having nice…
Density-based spatial clustering of applications with noise (DBSCAN) is a data clustering algorithm which has the high-performance rate for dataset where clusters have the constant density of data points. One of the significant attributes…
We introduce an accurate and efficient method for a class of nonlocal potential evaluations with free boundary condition, including the 3D/2D Coulomb, 2D Poisson and 3D dipolar potentials. Our method is based on a Gaussian-sum approximation…
Data analysis in high-dimensional spaces aims at obtaining a synthetic description of a data set, revealing its main structure and its salient features. We here introduce an approach providing this description in the form of a topography of…
Cluster analysis plays a crucial role in database mining, and one of the most widely used algorithms in this field is DBSCAN. However, DBSCAN has several limitations, such as difficulty in handling high-dimensional large-scale data,…
We use the presently observed number density of large X-ray clusters and linear mass power spectra to constrain the shape parameter ($\Gamma$), the spectral index ($n$), the amplitude of matter density perturbations on the scale of $8…