English
Related papers

Related papers: Analyzing Dataset Annotation Quality Management in…

200 papers

A major challenge in Natural Language Processing is obtaining annotated data for supervised learning. An option is the use of crowdsourcing platforms for data annotation. However, crowdsourcing introduces issues related to the annotator's…

Data is the engine of modern computer vision, which necessitates collecting large-scale datasets. This is expensive, and guaranteeing the quality of the labels is a major challenge. In this paper, we investigate efficient annotation…

Computer Vision and Pattern Recognition · Computer Science 2021-04-27 Yuan-Hong Liao , Amlan Kar , Sanja Fidler

The quality of the dataset is crucial for ensuring optimal performance and reliability of downstream task models. However, datasets often contain noisy data inadvertently included during the construction process. Numerous attempts have been…

Computation and Language · Computer Science 2024-09-25 Juhwan Choi , Jungmin Yun , Kyohoon Jin , YoungBin Kim

Well-annotated datasets, as shown in recent top studies, are becoming more important for researchers than ever before in supervised machine learning (ML). However, the dataset annotation process and its related human labor costs remain…

Computation and Language · Computer Science 2021-08-24 Haozhan Sun , Chenchen Xu , Hanna Suominen

This paper does not describe a novel method. Instead, it studies an essential foundation for reliable benchmarking and ultimately real-world application of AI-based image analysis: generating high-quality reference annotations. Previous…

Computer Vision and Pattern Recognition · Computer Science 2024-07-29 Tim Rädsch , Annika Reinke , Vivienn Weru , Minu D. Tizabi , Nicholas Heller , Fabian Isensee , Annette Kopp-Schneider , Lena Maier-Hein

The use of learning-based techniques to achieve automated software vulnerability detection has been of longstanding interest within the software security domain. These data-driven solutions are enabled by large software vulnerability…

Software Engineering · Computer Science 2023-01-16 Roland Croft , M. Ali Babar , Mehdi Kholoosi

Crowdsourcing platforms are often used to collect datasets for training machine learning models, despite higher levels of inaccurate labeling compared to expert labeling. There are two common strategies to manage the impact of such noise.…

Computation and Language · Computer Science 2022-06-14 Derek Chen , Zhou Yu , Samuel R. Bowman

Large amounts of annotated data have become more important than ever, especially since the rise of deep learning techniques. However, manual annotations are costly. We propose a tool that enables researchers to create large, high-quality,…

Digital Libraries · Computer Science 2021-12-23 Franziska Weeber , Felix Hamborg , Karsten Donnay , Bela Gipp

Data annotation is essential but highly error-prone in the development of AI-enabled perception systems (AIePS) for automated driving, and its quality directly influences model performance, safety, and reliability. However, the industry…

Software Engineering · Computer Science 2025-11-21 Hina Saeeda , Tommy Johansson , Mazen Mohamad , Eric Knauss

Annotation is the labeling of data by human effort. Annotation is critical to modern machine learning, and Bloomberg has developed years of experience of annotation at scale. This report captures a wealth of wisdom for applied annotation…

Computers and Society · Computer Science 2020-09-25 Tina Tseng , Amanda Stent , Domenic Maida

Language models have shown promise in various tasks but can be affected by undesired data during training, fine-tuning, or alignment. For example, if some unsafe conversations are wrongly annotated as safe ones, the model fine-tuned on…

Machine Learning · Computer Science 2024-03-26 Zhaowei Zhu , Jialu Wang , Hao Cheng , Yang Liu

The use of machine learning (ML)-based language models (LMs) to monitor content online is on the rise. For toxic text identification, task-specific fine-tuning of these models are performed using datasets labeled by annotators who provide…

Computation and Language · Computer Science 2021-12-08 Kofi Arhin , Ioana Baldini , Dennis Wei , Karthikeyan Natesan Ramamurthy , Moninder Singh

Identifying the quality of free-text arguments has become an important task in the rapidly expanding field of computational argumentation. In this work, we explore the challenging task of argument quality ranking. To this end, we created a…

Computation and Language · Computer Science 2019-11-27 Shai Gretz , Roni Friedman , Edo Cohen-Karlik , Assaf Toledo , Dan Lahav , Ranit Aharonov , Noam Slonim

Despite the high demand for manually annotated image data, managing complex and costly annotation projects remains under-discussed. This is partly due to the fact that leading such projects requires dealing with a set of diverse and…

Machine Learning · Computer Science 2025-08-21 Azim Ahmadzadeh , Rohan Adhyapak , Armin Iraji , Kartik Chaurasiya , V Aparna , Petrus C. Martens

The NLP community has long advocated for the construction of multi-annotator datasets to better capture the nuances of language interpretation, subjectivity, and ambiguity. This paper conducts a retrospective study to show how performance…

Computation and Language · Computer Science 2023-10-24 Pritam Kadasi , Mayank Singh

The increasing demand for high-quality datasets in machine learning has raised concerns about the ethical and responsible creation of these datasets. Dataset creators play a crucial role in developing responsible practices, yet their…

Machine Learning · Computer Science 2024-09-04 Will Orr , Kate Crawford

We explore the task of automatic assessment of argument quality. To that end, we actively collected 6.3k arguments, more than a factor of five compared to previously examined data. Each argument was explicitly and carefully annotated for…

Computation and Language · Computer Science 2019-09-04 Assaf Toledo , Shai Gretz , Edo Cohen-Karlik , Roni Friedman , Elad Venezian , Dan Lahav , Michal Jacovi , Ranit Aharonov , Noam Slonim

Current supervised deep learning frameworks rely on annotated data for modeling the underlying data distribution of a given task. In particular for computer vision algorithms powered by deep learning, the quality of annotated data is the…

Computer Vision and Pattern Recognition · Computer Science 2019-12-24 Joseph Nassar , Viveca Pavon-Harr , Marc Bosch , Ian McCulloh

Generative large language models (LLMs) can be a powerful tool for augmenting text annotation procedures, but their performance varies across annotation tasks due to prompt quality, text data idiosyncrasies, and conceptual difficulty.…

Computation and Language · Computer Science 2023-06-02 Nicholas Pangakis , Samuel Wolken , Neil Fasching

Datasets labelled by human annotators are widely used in the training and testing of machine learning models. In recent years, researchers are increasingly paying attention to label quality. However, it is not always possible to objectively…

Computer Vision and Pattern Recognition · Computer Science 2024-05-15 Luisa Schwirten , Jannes Scholz , Daniel Kondermann , Janis Keuper