English
Related papers

Related papers: If in a Crowdsourced Data Annotation Pipeline, a G…

200 papers

This paper explores using GPT-3.5 and GPT-4 to automate the data annotation process with automatic prompting techniques. The main aim of this paper is to reuse human annotation guidelines along with some annotated data to design automatic…

Computation and Language · Computer Science 2024-07-08 Sachin Yadav , Tejaswi Choppa , Dominik Schlechtweg

The rapid advancements in large language models (LLMs) have greatly expanded the potential for automated code-related tasks. Two primary methodologies are used in this domain: prompt engineering and fine-tuning. Prompt engineering involves…

Software Engineering · Computer Science 2025-02-21 Jiho Shin , Clark Tang , Tahmineh Mohati , Maleknaz Nayebi , Song Wang , Hadi Hemmati

Reference texts such as encyclopedias and news articles can manifest biased language when objective reporting is substituted by subjective writing. Existing methods to detect bias mostly rely on annotated data to train machine learning…

Computation and Language · Computer Science 2021-12-20 Timo Spinde , David Krieger , Manuel Plank , Bela Gipp

Crowdsourcing has been part of the IR toolbox as a cheap and fast mechanism to obtain labels for system development and evaluation. Successful deployment of crowdsourcing at scale involves adjusting many variables, a very important one…

Artificial Intelligence · Computer Science 2016-05-20 Ittai Abraham , Omar Alonso , Vasilis Kandylas , Rajesh Patel , Steven Shelford , Aleksandrs Slivkins

Crowdsourcing platforms provide marketplaces where task requesters can pay to get labels on their data. Such markets have emerged recently as popular venues for collecting annotations that are crucial in training machine learning models in…

Machine Learning · Computer Science 2017-08-28 Ashish Khetan , Sewoong Oh

This paper introduces a novel crowdsourcing worker selection algorithm, enhancing annotation quality and reducing costs. Unlike previous studies targeting simpler tasks, this study contends with the complexities of label interdependencies…

Computation and Language · Computer Science 2024-07-30 Yujie Wang , Chao Huang , Liner Yang , Zhixuan Fang , Yaping Huang , Yang Liu , Jingsi Yu , Erhong Yang

Crowdsourcing is a popular paradigm for effectively collecting labels at low cost. The Dawid-Skene estimator has been widely used for inferring the true labels from the noisy labels provided by non-expert crowdsourcing workers. However,…

Machine Learning · Statistics 2014-11-04 Yuchen Zhang , Xi Chen , Dengyong Zhou , Michael I. Jordan

Active learning algorithms automatically identify the most informative samples from large amounts of unlabeled data and tremendously reduce human annotation effort in inducing a machine learning model. In a conventional active learning…

Machine Learning · Computer Science 2026-04-28 Varun Totakura , Ankita Singh , Yushun Dong , Shayok Chakraborty

The ITU-T Recommendation P.808 provides a crowdsourcing approach for conducting a subjective assessment of speech quality using the Absolute Category Rating (ACR) method. We provide an open-source implementation of the ITU-T Rec. P.808 that…

Audio and Speech Processing · Electrical Eng. & Systems 2020-11-05 Babak Naderi , Ross Cutler

Crowdwork often entails tackling cognitively-demanding and time-consuming tasks. Crowdsourcing can be used for complex annotation tasks, from medical imaging to geospatial data, and such data powers sensitive applications, such as health…

Human-Computer Interaction · Computer Science 2020-09-07 Akira Matsui , Emilio Ferrara , Fred Morstatter , Andres Abeliuk , Aram Galstyan

Traditionally, psychophysical experiments are conducted by repeated measurements on a few well-trained participants under well-controlled conditions, often resulting in, if done properly, high quality data. In recent years, however,…

Machine Learning · Computer Science 2019-07-29 Siavash Haghiri , Patricia Rubisch , Robert Geirhos , Felix Wichmann , Ulrike von Luxburg

Human annotated data plays a crucial role in machine learning (ML) research and development. However, the ethical considerations around the processes and decisions that go into dataset annotation have not received nearly enough attention.…

Human-Computer Interaction · Computer Science 2022-06-22 Mark Diaz , Ian D. Kivlichan , Rachel Rosen , Dylan K. Baker , Razvan Amironesei , Vinodkumar Prabhakaran , Emily Denton

The vast collection of machine learning records available on the web presents a significant opportunity for meta-learning, where past experiments are leveraged to improve performance. Two crucial meta-learning tasks are pipeline performance…

Due to the difficulties in replicating and scaling up qualitative studies, such studies are rarely verified. Accordingly, in this paper, we leverage the advantages of crowdsourcing (low costs, fast speed, scalable workforce) to replicate…

Software Engineering · Computer Science 2017-03-03 Di Chen , Kathryn T. Stolee , Tim Menzies

In this study, a modular, data-free pipeline for multi-label intention recognition is proposed for agentic AI applications in transportation. Unlike traditional intent recognition systems that depend on large, annotated corpora and often…

Machine Learning · Computer Science 2025-11-06 Xiaocai Zhang , Hur Lim , Ke Wang , Zhe Xiao , Jing Wang , Kelvin Lee , Xiuju Fu , Zheng Qin

Crowdsourcing provides an efficient label collection schema for supervised machine learning. However, to control annotation cost, each instance in the crowdsourced data is typically annotated by a small number of annotators. This creates a…

Machine Learning · Computer Science 2021-07-23 Zhendong Chu , Hongning Wang

Disagreement in annotation is a common phenomenon in the development of NLP datasets and serves as a valuable source of insight. While majority voting remains the dominant strategy for aggregating labels, recent work has explored modeling…

We evaluated the capability of generative pre-trained transformers~(GPT-4) in analysis of textual data in tasks that require highly specialized domain expertise. Specifically, we focused on the task of analyzing court opinions to interpret…

Computation and Language · Computer Science 2023-10-05 Jaromir Savelka , Kevin D. Ashley , Morgan A Gray , Hannes Westermann , Huihui Xu

This paper proposes a method of abstractive summarization designed to scale to document collections instead of individual documents. Our approach applies a combination of semantic clustering, document size reduction within topic clusters,…

Artificial Intelligence · Computer Science 2023-10-10 Sengjie Liu , Christopher G. Healey

High-throughput phenotyping automates the mapping of patient signs to standardized ontology concepts and is essential for precision medicine. This study evaluates the automation of phenotyping of clinical summaries from the Online Mendelian…

Computation and Language · Computer Science 2025-06-11 Daniel B. Hier , S. Ilyas Munzir , Anne Stahlfeld , Tayo Obafemi-Ajayi , Michael D. Carrithers