English

Hierarchical Latent Word Clustering

Computation and Language 2016-01-22 v1

Abstract

This paper presents a new Bayesian non-parametric model by extending the usage of Hierarchical Dirichlet Allocation to extract tree structured word clusters from text data. The inference algorithm of the model collects words in a cluster if they share similar distribution over documents. In our experiments, we observed meaningful hierarchical structures on NIPS corpus and radiology reports collected from public repositories.

Keywords

Cite

@article{arxiv.1601.05472,
  title  = {Hierarchical Latent Word Clustering},
  author = {Halid Ziya Yerebakan and Fitsum Reda and Yiqiang Zhan and Yoshihisa Shinagawa},
  journal= {arXiv preprint arXiv:1601.05472},
  year   = {2016}
}
R2 v1 2026-06-22T12:33:48.959Z