Related papers: On a Clustering Criterion for Dependent Observatio…
Unsupervised machine learning, and in particular data clustering, is a powerful approach for the analysis of datasets and identification of characteristic features occurring throughout a dataset. It is gaining popularity across scientific…
Experimental evaluation is a major research methodology for investigating clustering algorithms and many other machine learning algorithms. For this purpose, a number of benchmark datasets have been widely used in the literature and their…
We study the problem of clustering sequences of unlabeled point sets taken from a common metric space. Such scenarios arise naturally in applications where a system or process is observed in distinct time intervals, such as biological…
Time series data may exhibit clustering over time and, in a multiple time series context, the clustering behavior may differ across the series. This paper is motivated by the Bayesian non--parametric modeling of the dependence between the…
We study a two-species bidirectional exclusion process, and a single species variant, which is motivated by the motion of organelles and vesicles along microtubules. Specifically, we are interested in the clustering of the particles and…
This paper presents a clustering approach that allows for rigorous statistical error control similar to a statistical test. We develop estimators for both the unknown number of clusters and the clusters themselves. The estimators depend on…
Many clustering methods, including k-means, require the user to specify the number of clusters as an input parameter. A variety of methods have been devised to choose the number of clusters automatically, but they often rely on strong…
This article proposes a mixture modeling approach to estimating cluster-wise conditional distributions in clustered (grouped) data. We adapt the mixture-of-experts model to the latent distributions, and propose a model in which each…
We present a general central limit theorem with simple, easy-to-check covariance-based sufficient conditions for triangular arrays of random vectors when all variables could be interdependent. The result is constructed from Stein's method,…
A data analysis pipeline is a structured sequence of steps that transforms raw data into meaningful insights by integrating multiple analysis algorithms. In many practical applications, analytical findings are obtained only after data pass…
Clustering is the process of finding underlying group structures in data. Although mixture model-based clustering is firmly established in the multivariate case, there is a relative paucity of work on matrix variate distributions and none…
For stationary sequences, under general local and asymptotic dependence restrictions, any limiting point process for time normalized upcrossings of high levels is a compound Poisson process, i.e., there is a clustering of high upcrossings,…
We establish a multivariate empirical process central limit theorem for stationary $\R^d$-valued stochastic processes $(X_i)_{i\geq 1}$ under very weak conditions concerning the dependence structure of the process. As an application we can…
Properties of strong mixing have been established for the stationary linear Hawkes process in the univariate case, and can serve as a basis for statistical applications. In this paper, we provide the technical arguments needed to extend the…
The Coupled-Cluster theory is one of the most successful high precision methods used to solve the stationary Schr\"odinger equation. In this article, we address the mathematical foundation of this theory with focus on the advances made in…
Fixed effects models are very flexible because they do not make assumptions on the distribution of effects and can also be used if the heterogeneity component is correlated with explanatory variables. A disadvantage is the large number of…
A computational theory for clustering and a semi-supervised clustering algorithm is presented. Clustering is defined to be the obtainment of groupings of data such that each group contains no anomalies with respect to a chosen grouping…
Under mild structural assumptions and regularity conditions on the marginal and conditional densities, an explicit bound on the $\beta$-mixing coefficients in terms of the physical dependence measure is provided. Consequently, weak physical…
Clustering evaluation measures are frequently used to evaluate the performance of algorithms. However, most measures are not properly normalized and ignore some information in the inherent structure of clusterings. We model the relation…
This paper considers functional central limit theorems for stationary absolutely regular mixing processes. Bounds for the entropy with bracketing are derived using recent results in Nickl and P\"otscher (2007). More specifically, their…