Related papers: Relative cluster entropy for power-law correlated …
The problem of adaptive noisy clustering is investigated. Given a set of noisy observations $Z_i=X_i+\epsilon_i$, $i=1,...,n$, the goal is to design clusters associated with the law of $X_i$'s, with unknown density $f$ with respect to the…
We present exact and asymptotic results for clusters in the one-dimensional totally asymmetric exclusion process (TASEP) with two different dynamics. The expected length of the largest cluster is shown to diverge logarithmically with…
In this note, we prove an $L^p$ uniform approximation of the fractional Brownian motion with Hurst exponent $0 < H < \frac{1}{2}$ by means of a family of continuous-time random walks imbedded on a given Brownian motion. The approximation is…
Relying on the excursion set theory, we compute the number density of local extrema and crossing statistics versus the threshold for the stock market indices. Comparing the number density of excursion sets calculated numerically with the…
Complex network theory crucially depends on the assumptions made about the degree distribution, while fitting degree distributions to network data is challenging, in particular for scale-free networks with power-law degrees. We present a…
Correlation clustering is arguably the most natural formulation of clustering. Given n objects and a pairwise similarity measure, the goal is to cluster the objects so that, to the best possible extent, similar objects are put in the same…
Power-law distributions are common, particularly in social physics. Here, we explore whether power-laws might arise as a consequence of a general variational principle for stochastic processes. We describe communities of 'social particles',…
Pragmatic trials evaluating health care interventions often adopt cluster randomization due to scientific or logistical considerations. Previous reviews have shown that co-primary endpoints are common in pragmatic trials but infrequently…
We study the Coulomb chain where particles are restricted to one dimension and experience three-dimensional Coulomb interactions with their nearest and next-to-nearest neighbours. The distances between consecutive particles are treated as…
Understanding the statistical laws governing citation dynamics remains a fundamental challenge in network theory and the science of science. Citation networks typically exhibit in-degree distributions well approximated by log-normal…
We consider an interacting particle system in continuous configuration space. The pair interaction has an attractive part. We show that, at low density, the system behaves approximately like an ideal mixture of clusters (droplets): we prove…
In this short note, we show how to use concentration inequalities in order to build exact confidence intervals for the Hurst parameter associated with a one-dimensional fractional Brownian motion
A new interpoint distance-based measure is proposed to identify the optimal number of clusters present in a data set. Designed in nonparametric approach, it is independent of the distribution of given data. Interpoint distances between the…
Comparing clusterings is central to evaluating unsupervised models, yet the many existing similarity measures can produce widely divergent, sometimes contradictory, evaluations. Clustering similarity measures are typically organized into…
The fidelity-based smooth min-relative entropy is a distinguishability measure that has appeared in a variety of contexts in prior work on quantum information, including resource theories like thermodynamics and coherence. Here we provide a…
Relative entropy serves as a cornerstone concept in quantum information theory. In this work, we study relative entropy of random states from major generic state models of Hilbert-Schmidt and Bures-Hall ensembles. In particular, we derive…
We develop Clustered Random Forests, a random forests algorithm for clustered data, arising from independent groups that exhibit within-cluster dependence. The leaf-wise predictions for each decision tree making up clustered random forests…
This paper studies clustering of data sequences using the k-medoids algorithm. All the data sequences are assumed to be generated from \emph{unknown} continuous distributions, which form clusters with each cluster containing a composite set…
Process discovery algorithms automatically extract process models from event logs, but high variability often results in complex and hard-to-understand models. To mitigate this issue, trace clustering techniques group process executions…
We develop an action formalism to calculate probabilities of rare events in cluster-cluster aggregation for arbitrary collision kernels and establish a pathwise large deviation principle with total mass being the rate. As an application,…