English

Easy, Reproducible and Quality-Controlled Data Collection with Crowdaq

Human-Computer Interaction 2020-10-15 v1

Abstract

High-quality and large-scale data are key to success for AI systems. However, large-scale data annotation efforts are often confronted with a set of common challenges: (1) designing a user-friendly annotation interface; (2) training enough annotators efficiently; and (3) reproducibility. To address these problems, we introduce Crowdaq, an open-source platform that standardizes the data collection pipeline with customizable user-interface components, automated annotator qualification, and saved pipelines in a re-usable format. We show that Crowdaq simplifies data annotation significantly on a diverse set of data collection use cases and we hope it will be a convenient tool for the community.

Keywords

Cite

@article{arxiv.2010.06694,
  title  = {Easy, Reproducible and Quality-Controlled Data Collection with Crowdaq},
  author = {Qiang Ning and Hao Wu and Pradeep Dasigi and Dheeru Dua and Matt Gardner and Robert L. Logan and Ana Marasovic and Zhen Nie},
  journal= {arXiv preprint arXiv:2010.06694},
  year   = {2020}
}

Comments

Accepted to the demo track of EMNLP 2020

R2 v1 2026-06-23T19:19:31.771Z