Is BERT Robust to Label Noise? A Study on Learning with Noisy Labels in Text Classification
Computation and Language
2022-04-21 v1
Abstract
Incorrect labels in training data occur when human annotators make mistakes or when the data is generated via weak or distant supervision. It has been shown that complex noise-handling techniques - by modeling, cleaning or filtering the noisy instances - are required to prevent models from fitting this label noise. However, we show in this work that, for text classification tasks with modern NLP models like BERT, over a variety of noise types, existing noisehandling methods do not always improve its performance, and may even deteriorate it, suggesting the need for further investigation. We also back our observations with a comprehensive analysis.
Cite
@article{arxiv.2204.09371,
title = {Is BERT Robust to Label Noise? A Study on Learning with Noisy Labels in Text Classification},
author = {Dawei Zhu and Michael A. Hedderich and Fangzhou Zhai and David Ifeoluwa Adelani and Dietrich Klakow},
journal= {arXiv preprint arXiv:2204.09371},
year = {2022}
}
Comments
Accepted at Workshop on Insights from Negative Results in NLP 2022 @ACL 2022