Evaluation of Representation Models for Text Classification with AutoML Tools
Computation and Language
2021-07-08 v2 Machine Learning
Abstract
Automated Machine Learning (AutoML) has gained increasing success on tabular data in recent years. However, processing unstructured data like text is a challenge and not widely supported by open-source AutoML tools. This work compares three manually created text representations and text embeddings automatically created by AutoML tools. Our benchmark includes four popular open-source AutoML tools and eight datasets for text classification purposes. The results show that straightforward text representations perform better than AutoML tools with automatically created text embeddings.
Cite
@article{arxiv.2106.12798,
title = {Evaluation of Representation Models for Text Classification with AutoML Tools},
author = {Sebastian Brändle and Marc Hanussek and Matthias Blohm and Maximilien Kintz},
journal= {arXiv preprint arXiv:2106.12798},
year = {2021}
}
Comments
Accecpted for Future Technologies Conference 2021