English

MixText: Linguistically-Informed Interpolation of Hidden Space for Semi-Supervised Text Classification

Computation and Language 2020-04-28 v1 Machine Learning

Abstract

This paper presents MixText, a semi-supervised learning method for text classification, which uses our newly designed data augmentation method called TMix. TMix creates a large amount of augmented training samples by interpolating text in hidden space. Moreover, we leverage recent advances in data augmentation to guess low-entropy labels for unlabeled data, hence making them as easy to use as labeled data.By mixing labeled, unlabeled and augmented data, MixText significantly outperformed current pre-trained and fined-tuned models and other state-of-the-art semi-supervised learning methods on several text classification benchmarks. The improvement is especially prominent when supervision is extremely limited. We have publicly released our code at https://github.com/GT-SALT/MixText.

Keywords

Cite

@article{arxiv.2004.12239,
  title  = {MixText: Linguistically-Informed Interpolation of Hidden Space for Semi-Supervised Text Classification},
  author = {Jiaao Chen and Zichao Yang and Diyi Yang},
  journal= {arXiv preprint arXiv:2004.12239},
  year   = {2020}
}

Comments

ACL 2020

R2 v1 2026-06-23T15:05:53.197Z