English
Related papers

Related papers: If in a Crowdsourced Data Annotation Pipeline, a G…

200 papers

HCI increasingly employs Machine Learning and Image Recognition, in particular for visual analysis of user interfaces (UIs). A popular way for obtaining human-labeled training data is Crowdsourcing, typically using the quality control…

Human-Computer Interaction · Computer Science 2020-12-29 Maxim Bakaev , Sebastian Heil , Martin Gaedke

This paper does not describe a novel method. Instead, it studies an essential foundation for reliable benchmarking and ultimately real-world application of AI-based image analysis: generating high-quality reference annotations. Previous…

Computer Vision and Pattern Recognition · Computer Science 2024-07-29 Tim Rädsch , Annika Reinke , Vivienn Weru , Minu D. Tizabi , Nicholas Heller , Fabian Isensee , Annette Kopp-Schneider , Lena Maier-Hein

Sentiment analysis of alternative tobacco products on social media is important for tobacco control research. Large Language Models (LLMs) can help streamline the labor-intensive human sentiment analysis process. This study examined the…

Computation and Language · Computer Science 2025-03-06 Kwanho Kim , Soojong Kim

With the success of modern internet based platform, such as Amazon Mechanical Turk, it is now normal to collect a large number of hand labeled samples from non-experts. The Dawid- Skene algorithm, which is based on Expectation- Maximization…

Machine Learning · Computer Science 2019-02-11 Changbo Zhu , Huan Xu , Shuicheng Yan

Supervised classification heavily depends on datasets annotated by humans. However, in subjective tasks such as toxicity classification, these annotations often exhibit low agreement among raters. Annotations have commonly been aggregated…

Computation and Language · Computer Science 2024-05-17 Negar Mokhberian , Myrl G. Marmarelis , Frederic R. Hopp , Valerio Basile , Fred Morstatter , Kristina Lerman

Software crowdsourcing platforms employ extrinsic rewards such as rating or ranking systems to motivate workers. Such rating systems are noisy and provide limited knowledge about workers' preferences and performance. To develop better…

Software Engineering · Computer Science 2022-01-20 Razieh Saremi , Hamid Shamszare , Marzieh Lotfalian Saremi , Ye Yang

We propose a streaming algorithm for the binary classification of data based on crowdsourcing. The algorithm learns the competence of each labeller by comparing her labels to those of other labellers on the same tasks and uses this…

Machine Learning · Statistics 2016-02-24 Thomas Bonald , Richard Combes

Large language models (LLMs) are increasingly capable and prevalent, and can be used to produce creative content. The quality of content is influenced by the prompt used, with more specific prompts that incorporate examples generally…

Computation and Language · Computer Science 2023-08-28 Samuel Rhys Cox , Ashraf Abdul , Wei Tsang Ooi

The growing need for labeled training data has made crowdsourcing an important part of machine learning. The quality of crowdsourced labels is, however, adversely affected by three factors: (1) the workers are not experts; (2) the…

Computer Science and Game Theory · Computer Science 2015-09-08 Nihar B. Shah , Dengyong Zhou , Yuval Peres

Study Objective: Machine learning models have advanced medical image processing and can yield faster, more accurate diagnoses. Despite a wealth of available medical imaging data, high-quality labeled data for model training is lacking. We…

Crowd-sourcing is a cheap and popular means of creating training and evaluation datasets for machine learning, however it poses the problem of `truth inference', as individual workers cannot be wholly trusted to provide reliable…

Machine Learning · Computer Science 2019-02-26 Yuan Li , Benjamin I. P. Rubinstein , Trevor Cohn

Annotating large datasets can be challenging. However, crowd-sourcing is often expensive and can lack quality, especially for non-trivial tasks. We propose a method of using LLMs as few-shot learners for annotating data in a complex natural…

Recent text generation research has increasingly focused on open-ended domains such as story and poetry generation. Because models built for such tasks are difficult to evaluate automatically, most researchers in the space justify their…

Computation and Language · Computer Science 2021-09-15 Marzena Karpinska , Nader Akoury , Mohit Iyyer

The phenomenal success of certain crowdsourced online platforms, such as Wikipedia, is accredited to their ability to tap the crowd's potential to collaboratively build knowledge. While it is well known that the crowd's collective wisdom…

Computers and Society · Computer Science 2015-08-28 Anamika Chhabra , S. R. S. Iyengar , Poonam Saini , Rajesh Shreedhar Bhat , Vijay Kumar

Crowdsourcing is widely used to create data for common natural language understanding tasks. Despite the importance of these datasets for measuring and refining model understanding of language, there has been little focus on the…

Computation and Language · Computer Science 2021-06-03 Nikita Nangia , Saku Sugawara , Harsh Trivedi , Alex Warstadt , Clara Vania , Samuel R. Bowman

Accurate and scalable annotation of medical data is critical for the development of medical AI, but obtaining time for annotation from medical experts is challenging. Gamified crowdsourcing has demonstrated potential for obtaining highly…

Typically crowdsourcing-based approaches to gather annotated data use inter-annotator agreement as a measure of quality. However, in many domains, there is ambiguity in the data, as well as a multitude of perspectives of the information…

Human-Computer Interaction · Computer Science 2018-08-21 Anca Dumitrache , Oana Inel , Lora Aroyo , Benjamin Timmermans , Chris Welty

Modern affective computing systems rely heavily on datasets with human-annotated emotion labels, for training and evaluation. However, human annotations are expensive to obtain, sensitive to study design, and difficult to quality control,…

Computation and Language · Computer Science 2024-12-12 Minxue Niu , Yara El-Tawil , Amrit Romana , Emily Mower Provost

Crowdsourcing requesters on Amazon Mechanical Turk (AMT) have raised questions about the reliability of the workers. The AMT workforce is very diverse and it is not possible to make blanket assumptions about them as a group. Some requesters…

Computation and Language · Computer Science 2021-11-10 Jessica Huynh , Jeffrey Bigham , Maxine Eskenazi

Employing multiple workers to label data for machine learning models has become increasingly important in recent years with greater demand to collect huge volumes of labelled data to train complex models while mitigating the risk of…

Artificial Intelligence · Computer Science 2021-02-18 Robert McCluskey , Amir Enshaei , Bashar Awwad Shiekh Hasan
‹ Prev 1 3 4 5 6 7 10 Next ›