English
Related papers

Related papers: A Bayesian latent allocation model for clustering …

200 papers

In this paper a relative number density parameter, called the neighborhood function, is introduced so that the crowded nature of the neighborhood of individual sources can be described. With this parameter one can determine the probability…

Astrophysics · Physics 2009-11-13 Yi-Ping Qin , Lian-Zhong Lv , Fu-Wen Zhang , Bin-Bin Zhang , Jin Zhang

Partitioning ocean flows into regions dynamically distinct from their surroundings based on material transport can assist search-and-rescue planning by reducing the search domain. The spectral clustering method partitions the domain by…

Atmospheric and Oceanic Physics · Physics 2020-08-28 Guilherme S. Vieira , Irina I. Rypina , Michael R. Allshouse

High-dimensional data clustering has become and remains a challenging task for modern statistics and machine learning, with a wide range of applications. We consider in this work the powerful discriminative latent mixture model, and we…

Methodology · Statistics 2020-12-09 Nicolas Jouvin , Charles Bouveyron , Pierre Latouche

Cluster number is typically a parameter selected at the outset in clustering problems, and while impactful, the choice can often be difficult to justify. Inspired by bioinformatics, this study examines how the nature of clusters varies with…

Machine Learning · Computer Science 2025-02-25 Justin Miller , Tristram Alexander

Various methods have been developed to combine inference across multiple sets of results for unsupervised clustering, within the ensemble clustering literature. The approach of reporting results from one `best' model out of several…

Bi-clustering is a useful approach in analyzing biological data when observations come from heterogeneous groups and have a large number of features. We outline a general Bayesian approach in tackling bi-clustering problems in moderate to…

Applications · Statistics 2021-02-11 Han Yan , Jiexing Wu , Yang Li , Jun S. Liu

For many taxonomic groups, online biodiversity portals used by naturalists and citizen scientists constitute the primary source of distributional information. Over the last decade, site-occupancy models have been advanced as a promising…

We begin by reviewing some probabilistic results about the Dirichlet Process and its close relatives, focussing on their implications for statistical modelling and analysis. We then introduce a class of simple mixture models in which…

Methodology · Statistics 2010-03-23 Peter J. Green

Modeling species abundance patterns using local environmental features is an important, current problem in ecology. The Cape Floristic Region (CFR) in South Africa is a global hot spot of diversity and endemism, and provides a rich class of…

Species distribution models usually attempt to explain presence-absence or abundance of a species at a site in terms of the environmental features (socalled abiotic features) present at the site. Historically, such models have considered…

Methodology · Statistics 2018-09-25 Shinichiro Shirota , Alan E. Gelfand , Sudipto Banerjee

Ecosystems tend to fluctuate around stable equilibria in response to internal dynamics and environmental factors. Occasionally, they enter an unstable tipping region and collapse into an alternative stable state. Our understanding of how…

We present a federated learning approach for Bayesian model-based clustering of large-scale binary and categorical datasets. We introduce a principled 'divide and conquer' inference procedure using variational inference with local merge and…

Machine Learning · Statistics 2025-11-13 Jackie Rao , Francesca L. Crowe , Tom Marshall , Sylvia Richardson , Paul D. W. Kirk

Cancer radiomics is an emerging discipline promising to elucidate lesion phenotypes and tumor heterogeneity through patterns of enhancement, texture, morphology, and shape. The prevailing technique for image texture analysis relies on the…

Applications · Statistics 2020-11-12 Xiao Li , Michele Guindani , Chaan S. Ng , Brian P. Hobbs

The problem of selecting a model given a set of candidates remains a challenging one that pervades many scientific fields. We employ techniques from the theory of Lie groups to analyse the symmetries in differential equation models of…

Quantitative Methods · Quantitative Biology 2023-10-10 Reemon Spector

To improve the predictability of complex computational models in the experimentally-unknown domains, we propose a Bayesian statistical machine learning framework utilizing the Dirichlet distribution that combines results of several…

Methodology · Statistics 2023-11-06 Vojtech Kejzlar , Léo Neufcourt , Witold Nazarewicz

The recent release of the second Gravitational-Wave Transient Catalog (GWTC-2) has increased significantly the number of known GW events, enabling unprecedented constraints on formation models of compact binaries. One pressing question is…

High Energy Astrophysical Phenomena · Physics 2021-04-28 Kaze W. K. Wong , Katelyn Breivik , Kyle Kremer , Thomas Callister

The last decades have not only been characterized by an explosive growth of data, but also an increasing appreciation of data as a valuable resource. Their value comes with the ability to extract meaningful patterns that are of economic,…

Machine Learning · Statistics 2020-02-27 Jonas I. Liechti , Sebastian Bonhoeffer

This paper proposes an early detection method for cluster structural changes. Cluster structure refers to discrete structural characteristics, such as the number of clusters, when data are represented using finite mixture models, such as…

Machine Learning · Statistics 2024-03-28 Kento Urano , Ryo Yuki , Kenji Yamanishi

Investigation of species abundance has become a vital component of many ecological monitoring studies. The primary objective of these studies is to understand how specific species are distributed across the study domain, as well as…

Applications · Statistics 2015-05-12 Guohui Wu , Scott H. Holan , Charles H. Nilon , Christopher K. Wikle

Biclustering is used for simultaneous clustering of the observations and variables when there is no group structure known \textit{a priori}. It is being increasingly used in bioinformatics, text analytics, etc. Previously, biclustering has…

Methodology · Statistics 2020-09-14 Wangshu Tu , Sanjeena Subedi