English

Clustering Network Tree Data From Respondent-driven sampling with application to opioid users in New York City

Social and Information Networks 2020-08-11 v1 Methodology

Abstract

There is great interest in finding meaningful subgroups of attributed network data. There are many available methods for clustering complete network. Unfortunately, much network data is collected through sampling, and therefore incomplete. Respondent-driven sampling (RDS) is a widely used method for sampling hard-to-reach human populations based on tracing links in the underlying unobserved social network. The resulting data therefore have tree structure representing a sub-sample of the network, along with many nodal attributes. In this paper, we introduce an approach to adjust mixture models for general network clustering for samplings by RDS. We apply our model to data on opioid users in New York City, and detect communities reflecting group characteristics of interest for intervention activities, including drug use patterns, social connections and other community variables

Keywords

Cite

@article{arxiv.2008.03604,
  title  = {Clustering Network Tree Data From Respondent-driven sampling with application to opioid users in New York City},
  author = {Shuaimin Kang and Krista Gile and Pedro Mateu-Gelabert and Honoria Guarino},
  journal= {arXiv preprint arXiv:2008.03604},
  year   = {2020}
}

Comments

25 pages

R2 v1 2026-06-23T17:43:33.889Z