English
Related papers

Related papers: Dirichlet-tree multinomial mixtures for clustering…

200 papers

Arguably the key issue in modelling discrete choice data is capturing preference heterogeneity. This can be through observed characteristics, and/or using techniques for capturing random heterogeneity across respondents. On the latter, in…

Methodology · Statistics 2025-06-18 Thomas O. Hancock , John Buckell

This paper explores test-agnostic long-tail recognition, a challenging long-tail task where the test label distributions are unknown and arbitrarily imbalanced. We argue that the variation in these distributions can be broken down…

Machine Learning · Computer Science 2026-03-02 Zhiyong Yang , Qianqian Xu , Sicong Li , Zitai Wang , Xiaochun Cao , Qingming Huang

Scientific studies in the last two decades have established the central role of the microbiome in disease and health. Differential abundance analysis seeks to identify microbial taxa associated with sample groups defined by a factor such as…

Methodology · Statistics 2023-12-29 Archie Sachdeva , Somnath Datta , Subharup Guha

Causal mediation analysis in cluster-randomized trials (CRTs) is essential for explaining how cluster-level interventions affect individual outcomes, yet it is complicated by interference, post-treatment confounding, and hierarchical…

Methodology · Statistics 2026-03-25 Yuki Ohnishi , Michael J. Daniels , Lei Yang , Fan Li

Determining the number of clusters is a fundamental issue in data clustering. Several algorithms have been proposed, including centroid-based algorithms using the Euclidean distance and model-based algorithms using a mixture of probability…

Machine Learning · Computer Science 2024-07-30 Ryosuke Motegi , Yoichi Seki

The dissimilarity mixture autoencoder (DMAE) is a neural network model for feature-based clustering that incorporates a flexible dissimilarity function and can be integrated into any kind of deep learning architecture. It internally…

Machine Learning · Computer Science 2021-07-16 Juan S. Lara , Fabio A. González

Determining phenotypes of diseases can have considerable benefits for in-hospital patient care and to drug development. The structure of high dimensional data sets such as electronic health records are often represented through an embedding…

A mixture of joint generalized hyperbolic distributions (MJGHD) is introduced for asymmetric clustering for high-dimensional data. The MJGHD approach takes into account the cluster-specific subspace, thereby limiting the number of…

Methodology · Statistics 2018-11-02 Yang Tang , Ryan P. Browne , Paul D. McNicholas

We describe a nonparametric topic model for labeled data. The model uses a mixture of random measures (MRM) as a base distribution of the Dirichlet process (DP) of the HDP framework, so we call it the DP-MRM. To model labeled data, we…

Machine Learning · Computer Science 2012-06-22 Dongwoo Kim , Suin Kim , Alice Oh

The conventional use of the Generalized Extreme Value (GEV) distribution to model block maxima may be inappropriate when extremes are actually structured into multiple heterogeneous groups. In this work, we propose a novel approach for…

Neutral models which assume ecological equivalence between species provide null models for community assembly. In Hubbell's Unified Neutral Theory of Biodiversity (UNTB), many local communities are connected to a single metacommunity…

Populations and Evolution · Quantitative Biology 2020-09-15 Keith Harris , Todd L Parsons , Umer Z Ijaz , Leo Lahti , Ian Holmes , Christopher Quince

Several epidemiological studies have provided evidence that long-term exposure to fine particulate matter (PM2.5) increases mortality risk. Furthermore, some population characteristics (e.g., age, race, and socioeconomic status) might play…

Methodology · Statistics 2023-11-01 Dafne Zorzetto , Falco J. Bargagli-Stoffi , Antonio Canale , Francesca Dominici

Model-based clustering is widely used for identifying and distinguishing types of diseases. However, modern biomedical data coming with high dimensions make it challenging to perform the model estimation in traditional cluster analysis. The…

Methodology · Statistics 2025-07-22 Kazeem Kareem , Fan Dai

We consider the problem of clustering data points in high dimensions, i.e. when the number of data points may be much smaller than the number of dimensions. Specifically, we consider a Gaussian mixture model (GMM) with non-spherical…

Statistics Theory · Mathematics 2014-06-10 Martin Azizyan , Aarti Singh , Larry Wasserman

We consider the problem of analyzing the heterogeneity of clustering distributions for multiple groups of observed data, each of which is indexed by a covariate value, and inferring global clusters arising from observations aggregated over…

Methodology · Statistics 2012-12-06 XuanLong Nguyen

Directional data require specialized probability models because of the non-Euclidean and periodic nature of their domain. When a directional variable is observed jointly with linear variables, modeling their dependence adds an additional…

Methodology · Statistics 2022-12-22 Tong Zou , Hal S. Stern

Clustering aims to group similar objects together while separating dissimilar ones apart. Thereafter, structures hidden in data can be identified to help understand data in an unsupervised manner. Traditional clustering methods such as…

Computer Vision and Pattern Recognition · Computer Science 2023-06-23 Jiawei Yao , Enbei Liu , Maham Rashid , Juhua Hu

Dirichlet process mixture model (DPMM) is a popular Bayesian nonparametric model. In this paper, we apply this model to weighted data and then estimate the un-weighted distribution from the corresponding weighted distribution using the…

Computation · Statistics 2018-12-12 Soghra Bohlourihajjar , Soleiman Khazaei

The findings of the 2023 AAPM Grand Challenge on Deep Generative Modeling for Learning Medical Image Statistics are reported in this Special Report. The goal of this challenge was to promote the development of deep generative models (DGMs)…

Image and Video Processing · Electrical Eng. & Systems 2024-05-06 Rucha Deshpande , Varun A. Kelkar , Dimitrios Gotsis , Prabhat Kc , Rongping Zeng , Kyle J. Myers , Frank J. Brooks , Mark A. Anastasio

The paper introduces the DIverse MultiPLEx Generalized Dot Product Graph (DIMPLE-GDPG) network model where all layers of the network have the same collection of nodes and follow the Generalized Dot Product Graph (GDPG) model. In addition,…

Methodology · Statistics 2023-03-27 Marianna Pensky , Yaxuan Wang
‹ Prev 1 8 9 10 Next ›