English
Related papers

Related papers: Frequency of Frequencies Distributions and Size De…

200 papers

Background: We study the statistical properties of fragment coverage in genome sequencing experiments. In an extension of the classic Lander-Waterman model, we consider the effect of the length distribution of fragments. We also introduce…

Genomics · Quantitative Biology 2010-05-03 Steven N. Evans , Valerie Hower , Lior Pachter

In this paper we study random partitions of 1,...n, where every cluster of size j can be in any of w\_j possible internal states. The Gibbs (n,k,w) distribution is obtained by sampling uniformly among such partitions with k clusters. We…

Probability · Mathematics 2007-05-23 Nathanael Berestycki , Jim Pitman

This paper presents and analyzes an approach to cluster-based inference for dependent data. The primary setting considered here is with spatially indexed data in which the dependence structure of observed random variables is characterized…

Statistics Theory · Mathematics 2022-11-16 Jianfei Cao , Christian Hansen , Damian Kozbur , Lucciano Villacorta

In this paper a relative number density parameter, called the neighborhood function, is introduced so that the crowded nature of the neighborhood of individual sources can be described. With this parameter one can determine the probability…

Astrophysics · Physics 2009-11-13 Yi-Ping Qin , Lian-Zhong Lv , Fu-Wen Zhang , Bin-Bin Zhang , Jin Zhang

This paper presents a new derivation of the Generalized Poisson distribution. This distribution provides a good fit to the evolved, counts-in-cells distribution measured in numerical simulations of hierarchical clustering from Poisson…

Astrophysics · Physics 2009-10-30 Ravi K. Sheth

When observations are organized into groups where commonalties exist amongst them, the dependent random measures can be an ideal choice for modeling. One of the propositions of the dependent random measures is that the atoms of the…

Machine Learning · Statistics 2016-06-28 Cheng Luo , Richard Yi Da Xu , Yang Xiang

In this paper, we consider the problem of partitioning a small data sample of size $n$ drawn from a mixture of $2$ sub-gaussian distributions. Our work is motivated by the application of clustering individuals according to their population…

Statistics Theory · Mathematics 2023-01-05 Shuheng Zhou

A cluster tree provides a highly-interpretable summary of a density function by representing the hierarchy of its high-density clusters. It is estimated using the empirical tree, which is the cluster tree constructed from a density…

Statistics Theory · Mathematics 2017-02-14 Jisu Kim , Yen-Chi Chen , Sivaraman Balakrishnan , Alessandro Rinaldo , Larry Wasserman

The inclusion of a fragmentation mechanism in population balance equations introduces complex interactions that make the analytical or even computational treatment much more difficult than for the pure aggregation case. This is specially…

Disordered Systems and Neural Networks · Physics 2022-03-14 Arturo Berrones-Santos , Luis Benavides-Vázquez , Elisa Schaeffer , Javier Almaguer

We propose a general model of unweighted and undirected networks having the scale-free property and fractal nature. Unlike the existing models of fractal scale-free networks (FSFNs), the present model can systematically and widely change…

Physics and Society · Physics 2022-05-04 Kousuke Yakubo , Yuka Fujiki

Scale invariance (fractality) is a prominent feature of the large-scale behavior of many stochastic systems. In this work, we construct an algorithm for the statistical identification of the Hurst distribution (in particular, the scaling…

Methodology · Statistics 2025-01-31 Patrice Abry , Gustavo Didier , Oliver Orejola , Herwig Wendt

Categorical random variables are a common staple in machine learning methods and other applications across disciplines. Many times, correlation within categorical predictors exists, and has been noted to have an effect on various algorithm…

Probability · Mathematics 2017-01-25 Rachel Traylor

We consider a weighted sum of a series of independent Poisson random variables and show that it results in a new compound Poisson distribution which includes the Poisson distribution and Poisson distribution of order k. An explicit…

Probability · Mathematics 2025-06-18 Palaniappan Vellaisamy , Tomoyuki Ichiba

By the method of Poissonization we confirm some existing results concerning consistent estimation of the structural distribution function in the situation of a large number of rare events. Inconsistency of the so called natural estimator is…

Statistics Theory · Mathematics 2007-06-13 Bert van Es , Stamatis Kolios

With inspiration from Random Forests (RF) in the context of classification, a new clustering ensemble method---Cluster Forests (CF) is proposed. Geometrically, CF randomly probes a high-dimensional data cloud to obtain "good local…

Methodology · Statistics 2013-06-07 Donghui Yan , Aiyou Chen , Michael I. Jordan

Model-based clustering is a powerful tool that is often used to discover hidden structure in data by grouping observational units that exhibit similar response values. Recently, clustering methods have been developed that permit…

Methodology · Statistics 2025-06-24 Sally Paganin , Garritt L. Page , Fernando Andrés Quintana

The distribution of age-ordered frequencies arising from an exchangeable Gibbs partition is studied in relation with the distribution of the positions at which new mutations appear in a sample.

Probability · Mathematics 2007-07-10 Robert C. Griffiths , Dario Spanó

We study the cluster size distribution of particles for a two-species exclusion process which involves totally asymmetric transport process of two oppositely directed species with stochastic directional switching of the species on a 1D…

Statistical Mechanics · Physics 2022-09-13 Jim Chacko , Sudipto Muhuri , Goutam Tripathy

Random discrete distributions, say $F,$ known as species sampling models, represent a rich class of models for classification and clustering, in Bayesian statistics and machine learning. They also arise in various areas of probability and…

Statistics Theory · Mathematics 2019-08-21 Lanelot F. James

Probabilistic networks display a wide range of high average clustering coefficients independent of the number of nodes in the network. In particular, the local clustering coefficient decreases with the degree of the subtending node in a…

Physics and Society · Physics 2013-11-26 Vijay K Samalam