Related papers: Test of multiscaling in DLA model using an off-lat…
We consider the problem of clustering data points in high dimensions, i.e. when the number of data points may be much smaller than the number of dimensions. Specifically, we consider a Gaussian mixture model (GMM) with non-spherical…
Predicting urban growth is important for practical reasons, and also for the challenge it presents to theoretical frameworks for cluster dynamics. Recently, the model of diffusion limited aggregation (DLA) has been applied to describe urban…
Classical multidimensional scaling is an important dimension reduction technique. Yet few theoretical results characterizing its statistical performance exist. This paper provides a theoretical framework for analyzing the quality of…
The problem of dimension reduction is of increasing importance in modern data analysis. In this paper, we consider modeling the collection of points in a high dimensional space as a union of low dimensional subspaces. In particular we…
High-dimensional datasets often contain multiple meaningful clusterings in different subspaces. For example, objects can be clustered either by color, weight, or size, revealing different interpretations of the given dataset. A variety of…
We study driven particle systems with excluded volume interactions on a two-lane ladder with periodic boundaries, using Monte Carlo simulation, cluster mean-field theory, and numerical solution of the master equation. Particles in one lane…
Fractals represent one of the fundamental manifestations of complexity, and fractal networks serve as tools for characterizing and investigating the fractal structures and properties of large-scale systems. Higher-order networks have…
Functional data analysis (FDA) methods have computational and theoretical appeals for some high dimensional data, but lack the scalability to modern large sample datasets. To tackle the challenge, we develop randomized algorithms for two…
We consider internal diffusion limited aggregation in dimension larger than or equal to two. This is a random cluster growth model, where random walks start at the origin of the d-dimensional lattice, one at a time, and stop moving when…
Spectral clustering is a celebrated algorithm that partitions objects based on pairwise similarity information. While this approach has been successfully applied to a variety of domains, it comes with limitations. The reason is that there…
We propose a new class of models for variable clustering called Asymptotic Independent block (AI-block) models, which defines population-level clusters based on the independence of the maxima of a multivariate stationary mixing random…
Overlapping clusters are common in models of many practical data-segmentation applications. Suppose we are given $n$ elements to be clustered into $k$ possibly overlapping clusters, and an oracle that can interactively answer queries of the…
Diffusion on a T fractal lattice under the influence of topological biasing fields is studied by finite size scaling methods. This allows to avoid proliferation and singularities which would arise in a renormalization group approach on…
We show that the fractal growth described by the dielectric breakdown model exhibits a phase transition in the multifractal spectrum of the growth measure. The transition takes place because the tip-splitting of branches forms a fixed…
We present a structural clustering algorithm for large-scale datasets of small labeled graphs, utilizing a frequent subgraph sampling strategy. A set of representatives provides an intuitive description of each cluster, supports the…
Order parameter fluctuations (the largest cluster size distribution) are studied within a three-dimensional bond percolation model on small lattices. Cumulant ratios measuring the fluctuations exhibit distinct features near the percolation…
Disobeying the classical wisdom of statistical learning theory, modern deep neural networks generalize well even though they typically contain millions of parameters. Recently, it has been shown that the trajectories of iterative…
This paper studies the large-scale subspace clustering (LSSC) problem with million data points. Many popular subspace clustering methods cannot directly handle the LSSC problem although they have been considered as state-of-the-art methods…
Many approaches to 3D image segmentation are based on hierarchical clustering of supervoxels into image regions. Here we describe a distributed algorithm capable of handling a tremendous number of supervoxels. The algorithm works…
In binary classification, imbalance refers to situations in which one class is heavily under-represented. This issue is due to either a data collection process or because one class is indeed rare in a population. Imbalanced classification…