Text Augmentations with R-drop for Classification of Tweets Self Reporting Covid-19
Computation and Language
2023-11-08 v1 Information Retrieval
Machine Learning
Abstract
This paper presents models created for the Social Media Mining for Health 2023 shared task. Our team addressed the first task, classifying tweets that self-report Covid-19 diagnosis. Our approach involves a classification model that incorporates diverse textual augmentations and utilizes R-drop to augment data and mitigate overfitting, boosting model efficacy. Our leading model, enhanced with R-drop and augmentations like synonym substitution, reserved words, and back translations, outperforms the task mean and median scores. Our system achieves an impressive F1 score of 0.877 on the test set.
Cite
@article{arxiv.2311.03420,
title = {Text Augmentations with R-drop for Classification of Tweets Self Reporting Covid-19},
author = {Sumam Francis and Marie-Francine Moens},
journal= {arXiv preprint arXiv:2311.03420},
year = {2023}
}
Comments
This paper has been peer-reviewed and accepted for presentation at SMM4H'23 at AMIA 2023 Annual Symposium