English

Corpus Statistics in Text Classification of Online Data

Computation and Language 2018-03-20 v1 Information Retrieval Machine Learning

Abstract

Transformation of Machine Learning (ML) from a boutique science to a generally accepted technology has increased importance of reproduction and transportability of ML studies. In the current work, we investigate how corpus characteristics of textual data sets correspond to text classification results. We work with two data sets gathered from sub-forums of an online health-related forum. Our empirical results are obtained for a multi-class sentiment analysis application.

Keywords

Cite

@article{arxiv.1803.06390,
  title  = {Corpus Statistics in Text Classification of Online Data},
  author = {Marina Sokolova and Victoria Bobicev},
  journal= {arXiv preprint arXiv:1803.06390},
  year   = {2018}
}

Comments

12 pages, 6 tables, 1 figure

R2 v1 2026-06-23T00:55:54.963Z