English

The Parrot Dilemma: Human-Labeled vs. LLM-augmented Data in Classification Tasks

Computation and Language 2024-02-06 v2 Computers and Society Physics and Society

Abstract

In the realm of Computational Social Science (CSS), practitioners often navigate complex, low-resource domains and face the costly and time-intensive challenges of acquiring and annotating data. We aim to establish a set of guidelines to address such challenges, comparing the use of human-labeled data with synthetically generated data from GPT-4 and Llama-2 in ten distinct CSS classification tasks of varying complexity. Additionally, we examine the impact of training data sizes on performance. Our findings reveal that models trained on human-labeled data consistently exhibit superior or comparable performance compared to their synthetically augmented counterparts. Nevertheless, synthetic augmentation proves beneficial, particularly in improving performance on rare classes within multi-class tasks. Furthermore, we leverage GPT-4 and Llama-2 for zero-shot classification and find that, while they generally display strong performance, they often fall short when compared to specialized classifiers trained on moderately sized training sets.

Keywords

Cite

@article{arxiv.2304.13861,
  title  = {The Parrot Dilemma: Human-Labeled vs. LLM-augmented Data in Classification Tasks},
  author = {Anders Giovanni Møller and Jacob Aarup Dalsgaard and Arianna Pera and Luca Maria Aiello},
  journal= {arXiv preprint arXiv:2304.13861},
  year   = {2024}
}

Comments

Accepted at EACL 2024. 14 pages, 4 figures, 2 tables

R2 v1 2026-06-28T10:19:08.750Z