English

A Comparison of Word-based and Context-based Representations for Classification Problems in Health Informatics

Computation and Language 2019-06-14 v1 Information Retrieval

Abstract

Distributed representations of text can be used as features when training a statistical classifier. These representations may be created as a composition of word vectors or as context-based sentence vectors. We compare the two kinds of representations (word versus context) for three classification problems: influenza infection classification, drug usage classification and personal health mention classification. For statistical classifiers trained for each of these problems, context-based representations based on ELMo, Universal Sentence Encoder, Neural-Net Language Model and FLAIR are better than Word2Vec, GloVe and the two adapted using the MESH ontology. There is an improvement of 2-4% in the accuracy when these context-based representations are used instead of word-based representations.

Keywords

Cite

@article{arxiv.1906.05468,
  title  = {A Comparison of Word-based and Context-based Representations for Classification Problems in Health Informatics},
  author = {Aditya Joshi and Sarvnaz Karimi and Ross Sparks and Cecile Paris and C Raina MacIntyre},
  journal= {arXiv preprint arXiv:1906.05468},
  year   = {2019}
}

Comments

To Appear in the 18th ACL Workshop on Biomedical Natural Language Processing (BioNLP)

R2 v1 2026-06-23T09:52:16.743Z