Related papers: Leave-one-out Singular Subspace Perturbation Analy…
We investigate a clustering problem with data from a mixture of Gaussians that share a common but unknown, and potentially ill-conditioned, covariance matrix. We start by considering Gaussian mixtures with two equally-sized components and…
Biclustering is used for simultaneous clustering of the observations and variables when there is no group structure known \textit{a priori}. It is being increasingly used in bioinformatics, text analytics, etc. Previously, biclustering has…
Clustering, or unsupervised classification, is a task often plagued by outliers. Yet there is a paucity of work on handling outliers in clustering. Outlier identification algorithms tend to fall into three broad categories: outlier…
Mixture distributions arise in many parametric and non-parametric settings -- for example, in Gaussian mixture models and in non-parametric estimation. It is often necessary to compute the entropy of a mixture, but, in most cases, this…
The variation of spectral subspaces for linear self-adjoint operators under an additive bounded perturbation is considered. The objective is to estimate the norm of the difference of two spectral projections associated with isolated parts…
Spectral clustering is one of the most popular clustering methods for finding clusters in a graph, which has found many applications in data mining. However, the input graph in those applications may have many missing edges due to error in…
High-order clustering aims to classify objects in multiway datasets that are prevalent in various fields such as bioinformatics, recommendation systems, and social network analysis. Such data are often sparse and high-dimensional, posing…
The random-cluster model has been widely studied as a unifying framework for random graphs, spin systems and electrical networks, but its dynamics have so far largely resisted analysis. In this paper we analyze the Glauber dynamics of the…
We analyze the performance of spectral clustering for community extraction in stochastic block models. We show that, under mild conditions, spectral clustering applied to the adjacency matrix of the network can consistently recover hidden…
Crossing symmetry asserts that particles are indistinguishable from anti-particles traveling back in time. In quantum field theory, this statement translates to the long-standing conjecture that probabilities for observing the two scenarios…
In the paper [25], written in collaboration with Gesine Reinert, we proved a universality principle for the Gaussian Wiener chaos. In the present work, we aim at providing an original example of application of this principle in the…
Clustering is the process of finding underlying group structures in data. Although mixture model-based clustering is firmly established in the multivariate case, there is a relative paucity of work on matrix variate distributions and none…
In this paper, a scale mixture of Normal distributions model is developed for classification and clustering of data having outliers and missing values. The classification method, based on a mixture model, focuses on the introduction of…
We consider nearest neighbor spacing distributions of composite ensembles of levels. These are obtained by combining independently unfolded sequences of levels containing only few levels each. Two problems arise in the spectral analysis of…
Since the appearance of the classical paper of Lifshitz almost half a century ago, linear stability analysis of cosmological models is textbook knowledge. Until recently, however, little was known about the behavior of higher than linear…
This paper establishes a variant of Stewart's theorem (Theorem~6.4 of Stewart, {\em SIAM Rev.}, 15:727--764, 1973) for the singular subspaces associated with the SVD of a matrix subject to perturbations. Stewart's original version uses both…
Anomalous coarsening in far-from equilibrium one-dimensional systems is investigated by simulation and analytic techniques. The minimal hard core particle (exclusion) models contain mechanisms of aggregated particle diffusion, with rates…
We consider the problem of clustering noisy high-dimensional data points into a union of low-dimensional subspaces and a set of outliers. The number of subspaces, their dimensions, and their orientations are unknown. A probabilistic…
We study the problem of clustering $T$ trajectories of length $H$, each generated by one of K unknown ergodic Markov chains over a finite state space of size $S$. We derive an instance-dependent, high-probability lower bound on the…
We propose to study unitary matrix ensembles defined in terms of unitary stochastic transition matrices associated with Markov processes on graphs. We argue that the spectral statistics of such an ensemble (after ensemble averaging) depends…