English
Related papers

Related papers: Clustering South African households based on their…

200 papers

Sequence analysis is an increasingly popular approach for analysing life courses represented by ordered collections of activities experienced by subjects over time. Here, we analyse a survey data set containing information on the career…

In low- and middle-income countries, household surveys are the most reliable data source to examine health and demographic indicators at the subnational level, an exercise in small area estimation. Model-based unit-level models are favored…

Methodology · Statistics 2022-09-23 Yunhan Wu , Jon Wakefield

We introduce a novel statistical significance-based approach for clustering hierarchical data using semi-parametric linear mixed-effects models designed for responses with laws in the exponential family (e.g., Poisson and Bernoulli). Within…

Methodology · Statistics 2025-02-04 Alessandra Ragni , Chiara Masci , Francesca Ieva , Anna Maria Paganoni

In model-based clustering and classification, the cluster-weighted model constitutes a convenient approach when the random vector of interest constitutes a response variable Y and a set p of explanatory variables X. However, its…

Methodology · Statistics 2013-07-23 Sanjeena Subedi , Antonio Punzo , Salvatore Ingrassia , Paul D. McNicholas

Multivariate longitudinal data of mixed-type are increasingly collected in many science domains. However, algorithms to cluster this kind of data remain scarce, due to the challenge to simultaneously model the within- and between-time…

Machine Learning · Statistics 2025-09-16 Francesco Amato , Julien Jacques

We consider the estimation of wealth inequality measures with their confidence interval, based on survey data with interval censoring. We rely on a Bayesian hierarchical model. It consists of a model where, due to survey sampling and unit…

Applications · Statistics 2011-08-01 Eric Gautier

The LIPGENE-SU.VI.MAX study, like many others, recorded high dimensional continuous phenotypic data and categorical genotypic data. LIPGENE-SU.VI.MAX focuses on the need to account for both phenotypic and genetic factors when studying the…

The present study uses domain experts to estimate welfare levels and indicators from high-resolution satellite imagery. We use the wealth quintiles from the 2015 Tanzania DHS dataset as ground truth data. We analyse the performance of the…

General Economics · Economics 2022-10-18 Wahab Ibrahim , Ola Hall

We are concerned in clustering continuous data sets subject to non-ignorable missingness. We perform clustering with a specific semi-parametric mixture, under the assumption of conditional independence given the component. The mixture model…

Methodology · Statistics 2021-07-20 Marie Du Roy de Chaumaray , Matthieu Marbac

We propose a Bayesian approach for model-based clustering of multivariate categorical data where variables are allowed to be associated within clusters and the number of clusters is unknown. The approach uses a two-layer mixture of finite…

Methodology · Statistics 2024-07-09 Gertraud Malsiner-Walli , Bettina Grün , Sylvia Frühwirth-Schnatter

The study of international relations by definition deals with interdependencies among countries. One form of interdependence between countries is the diffusion of country-level features, such as policies, political regimes, or conflict. In…

Applications · Statistics 2019-03-22 Johan A. Elkink , Thomas U. Grund

In the field of population health research, understanding the similarities between geographical areas and quantifying their shared effects on health outcomes is crucial. In this paper, we synthesise a number of existing methods to create a…

Applications · Statistics 2023-11-27 Wala Draidi Areed , Aiden Price , Helen Thompson , Reid Malseed , Kerrie Mengersen

In low- and middle-income countries (LMICs), accurate estimates of subnational health and demographic indicators are critical for guiding policy and identifying disparities. Many indicators of interest are proportions of binary outcomes and…

Applications · Statistics 2025-06-23 Jon Wakefield , Peter A. Gao , Geir-Arne Fuglstad , Zehang Richard Li

Accurate, fine-grained poverty maps remain scarce across much of the Global South. While Demographic and Health Surveys (DHS) provide high-quality socioeconomic data, their spatial coverage is limited and reported coordinates are randomly…

Machine Learning · Computer Science 2025-11-04 Markus B. Pettersson , Adel Daoud

Synthetic indices are used in Economics to measure various aspects of monetary inequalities. These scalar indices take as input the distribution over a finite population, for example the population of a specific country. In this article we…

Applications · Statistics 2011-09-06 Eric Gautier

This note introduces an unsupervised learning algorithm to debug errors in finite element (FE) simulation models and details how it was productionised. The algorithm clusters degrees of freedom in the FE model using numerical properties of…

Computational Engineering, Finance, and Science · Computer Science 2023-10-26 Ramaseshan Kannan

Recent work on overfitting Bayesian mixtures of distributions offers a powerful framework for clustering multivariate data using a latent Gaussian model which resembles the factor analysis model. The flexibility provided by overfitting…

Methodology · Statistics 2019-08-29 Panagiotis Papastamoulis

Microbiome research has immense potential for unlocking insights into human health and disease. A common goal in human microbiome research is identifying subgroups of individuals with similar microbial composition that may be linked to…

Methodology · Statistics 2025-08-21 Suppapat Korsurat , Matthew D. Koslovsky

In some contexts, mixture models can fit certain variables well at the expense of others in ways beyond the analyst's control. For example, when the data include some variables with non-trivial amounts of missing values, the mixture model…

Methodology · Statistics 2016-09-06 Maria DeYoreo , Jerome P. Reiter , D. Sunshine Hillygus

In the framework of Bayesian model-based clustering based on a finite mixture of Gaussian distributions, we present a joint approach to estimate the number of mixture components and identify cluster-relevant variables simultaneously as well…

Methodology · Statistics 2016-06-23 Gertraud Malsiner-Walli , Sylvia Frühwirth-Schnatter , Bettina Grün