English
Related papers

Related papers: Sample Size Dependent Species Models

200 papers

By the method of Poissonization we confirm some existing results concerning consistent estimation of the structural distribution function in the situation of a large number of rare events. Inconsistency of the so called natural estimator is…

Statistics Theory · Mathematics 2007-06-13 Bert van Es , Stamatis Kolios

We study a two-species bidirectional exclusion process, and a single species variant, which is motivated by the motion of organelles and vesicles along microtubules. Specifically, we are interested in the clustering of the particles and…

Statistical Mechanics · Physics 2020-03-17 Jim Chacko , Sudipto Muhuri , Goutam Tripathy

Classical peaks over threshold analysis is widely used for statistical modeling of sample extremes, and can be supplemented by a model for the sizes of clusters of exceedances. Under mild conditions a compound Poisson process model allows…

Applications · Statistics 2016-08-14 Mária Süveges , Anthony C. Davison

Genome wide comparisons between enteric bacteria yield large sets of conserved putative regulatory sites on a gene by gene basis that need to be clustered into regulons. Using the assumption that regulatory sites can be represented as…

Biological Physics · Physics 2009-11-07 Erik van Nimwegen , Mihaela Zavolan , Nikolaus Rajewsky , Eric D. Siggia

We review the class of species sampling models (SSM). In particular, we investigate the relation between the exchangeable partition probability function (EPPF) and the predictive probability function (PPF). It is straightforward to define a…

Methodology · Statistics 2013-06-12 Jaeyong Lee , Fernando A. Quintana , Peter Müller , Lorenzo Trippa

We consider the task of modeling a dependent sequence of random partitions. It is well-known that a random measure in Bayesian nonparametrics induces a distribution over random partitions. The community has therefore assumed that the best…

Methodology · Statistics 2021-08-03 Garritt L. Page , Fernando A. Quintana , David B. Dahl

The problem of time-series clustering is considered in the case where each data-point is a sample generated by a piecewise stationary ergodic process. Stationary processes are perhaps the most general class of processes considered in…

Machine Learning · Statistics 2019-06-27 Azadeh Khaleghi , Daniil Ryabko

We consider a discrete latent variable model for two-way data arrays, which allows one to simultaneously produce clusters along one of the data dimensions (e.g. exchangeable observational units or features) and contiguous groups, or…

We propose a novel method for multiple clustering that assumes a co-clustering structure (partitions in both rows and columns of the data matrix) in each view. The new method is applicable to high-dimensional data. It is based on a…

An unsupervised classification method for point events occurring on a network of lines is proposed. The idea relies on the distributional flexibility and practicality of random partition models to discover the clustering structure featuring…

The distribution function of particles over clusters is proposed for a system of identical intersecting spheres, the centres of which are uniformly distributed in space. Consideration is based on the concept of the rank number of clusters,…

Statistical Mechanics · Physics 2023-02-01 Murat Kh. Khokonov , Azamat Kh. Khokonov

We propose the Plaid Atoms Model (PAM), a novel Bayesian nonparametric model for grouped data. Founded on an idea of `atom skipping', PAM is part of a well-established category of models that generate dependent random distributions and…

Methodology · Statistics 2024-01-02 Dehua Bi , Yuan Ji

This paper describes a compound Poisson-based random effects structure for modeling zero-inflated data. Data with large proportion of zeros are found in many fields of applied statistics, for example in ecology when trying to model and…

Applications · Statistics 2009-07-29 Marie-Pierre Etienne , Eric Parent , Benoit Hugues , Bernier Jacques

Survey data are often collected under multistage sampling designs where units are binned to clusters that are sampled in a first stage. The unit-indexed population variables of interest are typically dependent within cluster. We propose a…

Methodology · Statistics 2021-08-26 Luis G. Leon-Novelo , Terrance D. Savitsky

We consider estimation of the structural distribution function of the cell probabilities of a multinomial sample in situations where the number of cells is large. We review the performance of the natural estimator, an estimator based on…

Statistics Theory · Mathematics 2007-06-13 B. van Es , C. A. J. Klaassen , R. M. Mnatsakanov

We introduce a random partition model for Bayesian nonparametric regression. The model is based on infinitely-many disjoint regions of the range of a latent covariate-dependent Gaussian process. Given a realization of the process, the…

Methodology · Statistics 2013-01-04 George Karabatsos , Stephen G. Walker

An important functional of Poisson random measure is the negative binomial process (NBP). We use NBP to introduce a generalized Poisson-Kingman distribution and its corresponding random discrete probability measure. This random discrete…

Statistics Theory · Mathematics 2023-07-04 Sadegh Chegini , Mahmoud Zarepour

After generalizing the concept of clusters to incorporate clusters that are linked to other clusters through some relatively narrow bridges, an approach for detecting patches of separation between these clusters is developed based on an…

Computer Vision and Pattern Recognition · Computer Science 2020-01-10 Luciano da F. Costa

Dirichlet process mixtures are flexible non-parametric models, particularly suited to density estimation and probabilistic clustering. In this work we study the posterior distribution induced by Dirichlet process mixtures as the sample size…

Statistics Theory · Mathematics 2022-11-29 Filippo Ascolani , Antonio Lijoi , Giovanni Rebaudo , Giacomo Zanella

Cluster analysis is a popular unsupervised learning tool used in many disciplines to identify heterogeneous sub-populations within a sample. However, validating cluster analysis results and determining the number of clusters in a data set…

Machine Learning · Statistics 2024-04-26 Ali Turfah , Xiaoquan Wen