English

Evaluation of Representation Models for Text Classification with AutoML Tools

Computation and Language 2021-07-08 v2 Machine Learning

Abstract

Automated Machine Learning (AutoML) has gained increasing success on tabular data in recent years. However, processing unstructured data like text is a challenge and not widely supported by open-source AutoML tools. This work compares three manually created text representations and text embeddings automatically created by AutoML tools. Our benchmark includes four popular open-source AutoML tools and eight datasets for text classification purposes. The results show that straightforward text representations perform better than AutoML tools with automatically created text embeddings.

Keywords

Cite

@article{arxiv.2106.12798,
  title  = {Evaluation of Representation Models for Text Classification with AutoML Tools},
  author = {Sebastian Brändle and Marc Hanussek and Matthias Blohm and Maximilien Kintz},
  journal= {arXiv preprint arXiv:2106.12798},
  year   = {2021}
}

Comments

Accecpted for Future Technologies Conference 2021

R2 v1 2026-06-24T03:32:33.269Z