Related papers: Measures on contour, polymer or animal models. A p…
Motivated by the fundamental problem of measuring species diversity, this paper introduces the concept of a cluster structure to define an exchangeable cluster probability function that governs the joint distribution of a random count and…
Clustering is a central approach for unsupervised learning. After clustering is applied, the most fundamental analysis is to quantitatively compare clusterings. Such comparisons are crucial for the evaluation of clustering methods as well…
A method is discussed that allows combining sets of differential or inclusive measurements. It is assumed that at least one measurement was obtained with simultaneously fitting a set of nuisance parameters, representing sources of…
We prove estimates at infinity of convolutions $f^{n\star}$ and densities of the corresponding compound Poisson measures for a class of radial decreasing densities on $\mathbb{R}^d$, $d \geq 1$, which are not convolution equivalent.…
This paper proposes a novel, nonparametric, interpoint distance-based measure to investigate whether there exist any groups in a set of given data, and if so then, how many groups are prevailing in total. It is a cluster accuracy index…
We observe $n$ inhomogeneous Poisson processes with covariates and aim at estimating their intensities. We assume that the intensity of each Poisson process is of the form $s (\cdot, x)$ where $x$ is the covariate and where $s$ is an…
The random-cluster model, a correlated bond percolation model, unifies a range of important models of statistical mechanics in one description, including independent bond percolation, the Potts model and uniform spanning trees. By…
The evolution, regulation and sustenance of biological complexity is determined by protein-protein interaction network that is filled with dynamic events. Recent experimental evidences point out that clustering of proteins has a vital role…
A convergence criterion of cluster expansion is presented in the case of an abstract polymer system with general pair interactions (i.e. not necessarily hard core or repulsive). As a concrete example, the low temperature disordered phase of…
In the focus of our attention is the asymptotic properties of the sequence of convex hulls which arise as a result of a peeling procedure applied to the convex hull generated by a Poisson point process. Processes of the considered type are…
We study a local thinning $T_r$ that retains a point with probability $p(n_r)$, where $n_r$ counts neighbors within radius $r$. For Poisson input with spatially varying intensity, we obtain an exact intensity via a Poisson--mixture formula…
Ensemble methods are among the state-of-the-art predictive modeling approaches. Applied to modern big data, these methods often require a large number of sub-learners, where the complexity of each learner typically grows with the size of…
A novel non-parametric estimator of the correlation between grouped measurements of a quantity is proposed in the presence of noise. This work is primarily motivated by functional brain network construction from fMRI data, where brain…
Polymers are an effective test-bed for studying topological constraints in condensed matter due to a wide array of synthetically-available chain topologies. When linear and ring polymers are blended together, emergent rheological properties…
Probabilistic modeling provides the capability to represent and manipulate uncertainty in data, models, predictions and decisions. We are concerned with the problem of learning probabilistic models of dynamical systems from measured data.…
Mode clustering is a nonparametric method for clustering that defines clusters using the basins of attraction of a density estimator's modes. We provide several enhancements to mode clustering: (i) a soft variant of cluster assignment, (ii)…
This paper addresses the case where data come as point sets, or more generally as discrete measures. Our motivation is twofold: first we intend to approximate with a compactly supported measure the mean of the measure generating process,…
A new interpoint distance-based measure is proposed to identify the optimal number of clusters present in a data set. Designed in nonparametric approach, it is independent of the distribution of given data. Interpoint distances between the…
Peer-grouping is used in many sectors for organisational learning, policy implementation, and benchmarking. Clustering provides a statistical, data-driven method for constructing meaningful peer groups, but peer groups must be compatible…
This paper develops a theory of clustering and coding which combines a geometric model with a probabilistic model in a principled way. The geometric model is a Riemannian manifold with a Riemannian metric, ${g}_{ij}({\bf x})$, which we…