Related papers: Size-Distribution Scaling in Clusters of Allelomim…
Large language models with a huge number of parameters, when trained on near internet-sized number of tokens, have been empirically shown to obey neural scaling laws: specifically, their performance behaves predictably as a power law in…
The consistency of the maximum likelihood estimator for mixtures of elliptically-symmetric distributions for estimating its population version is shown, where the underlying distribution $P$ is nonparametric and does not necessarily belong…
Cluster analysis is a widely applied machine learning technique to understand the existing patterns in the population of gamma-ray bursts (GRBs), in order to explore their physical sources. In the present scenario, the number of clusters…
We investigate the random eigenvalues coming from the beta-Laguerre ensemble with parameter p, which is a generalization of the real, complex and quaternion Wishart matrices of parameter (n,p). In the case that the sample size n is much…
(abridged) We use a theoretical model to predict the clustering properties of galaxy clusters. Our technique accounts for past light-cone effects on the observed clustering and follows the non-linear evolution of the dark matter correlation…
The APM Cluster Survey was based on a modification of Abell's original classification scheme for galaxy clusters. Here we discuss the results of an investigation of the stability of the statistical properties of the cluster catalogue to…
We present generalized dynamical models describing the sharing of information, and the corresponding herd behavior, in a population based on the recent model proposed by Egu\'{\i}luz and Zimmermann (EZ) [Phys. Rev. Lett. 85, 5659 (2000)].…
The scaling properties of the cluster size distribution of a system of diffusing clusters is studied in terms of a simple kinetic mean field model. It is shown that a one parameter family of mathematically valid scaling solutions exists.…
Clustering, or grouping, dataset elements based on similarity can be used not only to classify a dataset into a few categories, but also to approximate it by a relatively large number of representative elements. In the latter scenario,…
Cluster sampling is common in survey practice, and the corresponding inference has been predominantly design-based. We develop a Bayesian framework for cluster sampling and account for the design effect in the outcome modeling. We consider…
The effect of introducing a mass dependent diffusion rate ~ m^{-alpha} in a model of coagulation with single-particle break up is studied both analytically and numerically. The model with alpha=0 is known to undergo a nonequilibrium phase…
We study the problem of graph clustering under a broad class of objectives in which the quality of a cluster is defined based on the ratio between the number of edges in the cluster, and the total weight of vertices in the cluster. We show…
A two parameter percolation model with nucleation and growth of finite clusters is developed taking the initial seed concentration \rho and a growth parameter g as two tunable parameters. Percolation transition is determined by the final…
We provide a complete asymptotic distribution theory for clustered data with a large number of independent groups, generalizing the classic laws of large numbers, uniform laws, central limit theory, and clustered covariance matrix…
Clustered data are common in practice. Clustering arises when subjects are measured repeatedly, or subjects are nested in groups (e.g., households, schools). It is often of interest to evaluate the correlation between two variables with…
We investigate the behavioral patterns of a population of agents, each controlled by a simple biologically motivated neural network model, when they are set in competition against each other in the Minority Model of Challet and Zhang. We…
Cluster-weighted factor analyzers (CWFA) are a versatile class of mixture models designed to estimate the joint distribution of a random vector that includes a response variable along with a set of explanatory variables. They are…
The evolution of the allelic proportion $x$ of a biallelic locus subject to the forces of mutation and drift is investigated in a diffusion model, assuming small scaled mutation rates. The overall scaled mutation rate is parametrized with…
Clustered standard errors and approximate randomization tests are popular inference methods that allow for dependence within observations. However, they require researchers to know the cluster structure ex ante. We propose a procedure to…
A model-based approach is developed for clustering categorical data with no natural ordering. The proposed method exploits the Hamming distance to define a family of probability mass functions to model the data. The elements of this family…