中文
相关论文

相关论文: Finding the Ground-Truth from Multiple Labellers: …

200 篇论文

Crowdsourcing systems aggregate decisions of many people to help users quickly identify high-quality options, such as the best answers to questions or interesting news stories. A long-standing issue in crowdsourcing is how option quality…

社会与信息网络 · 计算机科学 2020-10-28 Keith Burghardt , Tad Hogg , Raissa M. D'Souza , Kristina Lerman , Marton Posfai

We explore the design of an effective crowdsourcing system for an $M$-ary classification task. Crowd workers complete simple binary microtasks whose results are aggregated to give the final classification decision. We consider the scenario…

社会与信息网络 · 计算机科学 2017-04-05 Qunwei Li , Pramod K. Varshney

A growing literature on human-AI decision-making investigates strategies for combining human judgment with statistical models to improve decision-making. Research in this area often evaluates proposed improvements to models, interfaces, or…

计算机与社会 · 计算机科学 2023-05-29 Luke Guerdan , Amanda Coston , Zhiwei Steven Wu , Kenneth Holstein

Annotated images are required for both supervised model training and evaluation in image classification. Manually annotating images is arduous and expensive, especially for multi-labeled images. A recent trend for conducting such laboursome…

计算机视觉与模式识别 · 计算机科学 2022-12-07 Jianzhe Lin , Tianze Yu , Z. Jane Wang

Large language models often achieve strong benchmark gains without corresponding improvements in broader capability. We hypothesize that this discrepancy arises from differences in training regimes induced by data distribution. To…

机器学习 · 计算机科学 2026-04-10 Hongjian Zou , Yidan Wang , Qi Ding , Yixuan Liao , Xiaoxin Chen

How to better reduce measurement variability and bias introduced by subjectivity in crowdsourced labelling remains an open question. We introduce a theoretical framework for understanding how random error and measurement bias enter into…

人机交互 · 计算机科学 2023-12-05 Hasti Narimanzadeh , Arash Badie-Modiri , Iuliia Smirnova , Ted Hsuan Yun Chen

Recent trends in natural language processing research and annotation tasks affirm a paradigm shift from the traditional reliance on a single ground truth to a focus on individual perspectives, particularly in subjective tasks. In scenarios…

计算与语言 · 计算机科学 2024-04-18 Olufunke O. Sarumi , Béla Neuendorf , Joan Plepi , Lucie Flek , Jörg Schlötterer , Charles Welch

The questions in a crowdsourcing task typically exhibit varying degrees of difficulty and subjectivity. Their joint effects give rise to the variation in responses to the same question by different crowd-workers. This variation is low when…

人工智能 · 计算机科学 2018-02-15 Yuan Jin , Mark Carman , Ye Zhu , Wray Buntine

In machine learning, "ground truth" refers to the assumed correct labels used to train and evaluate models. However, the foundational "ground truth" paradigm rests on a positivistic fallacy that treats human disagreement as technical noise…

Distant supervision is a popular method for performing relation extraction from text that is known to produce noisy labels. Most progress in relation extraction and classification has been made with crowdsourced corrections to…

计算与语言 · 计算机科学 2022-09-21 Anca Dumitrache , Lora Aroyo , Chris Welty

Gathering training data is a key step of any supervised learning task, and it is both critical and expensive. Critical, because the quantity and quality of the training data has a high impact on the performance of the learned function.…

数据结构与算法 · 计算机科学 2021-10-28 Quentin Lutz , Élie de Panafieu , Alex Scott , Maya Stein

We present new methods for multilabel classification, relying on ensemble learning on a collection of random output graphs imposed on the multilabel and a kernel-based structured output learner as the base classifier. For ensemble learning,…

机器学习 · 计算机科学 2013-11-19 Hongyu Su , Juho Rousu

Many recent state-of-the-art results in language tasks were achieved using compound systems that perform multiple Language Model (LM) calls and aggregate their responses. However, there is little understanding of how the number of LM calls…

机器学习 · 计算机科学 2024-06-06 Lingjiao Chen , Jared Quincy Davis , Boris Hanin , Peter Bailis , Ion Stoica , Matei Zaharia , James Zou

The growing use of supervised machine learning in research and industry has increased the need for labeled datasets. Crowdsourcing has emerged as a popular method to create data labels. However, working on large batches of tasks leads to…

人机交互 · 计算机科学 2022-09-30 Chandramohan Sudar , Michael Froehlich , Florian Alt

Crowd-sourcing is an increasingly popular tool for image analysis in animal ecology. Computer vision methods that can utilize crowd-sourced annotations can help scale up analysis further. In this work we study the potential to do so on the…

计算机视觉与模式识别 · 计算机科学 2022-05-31 Justin Kay , Catherine M. Foley , Tom Hart

Distant supervision for relation extraction enables one to effectively acquire structured relations out of very large text corpora with less human efforts. Nevertheless, most of the prior-art models for such tasks assume that the given text…

计算与语言 · 计算机科学 2019-09-13 Junfan Chen , Richong Zhang , Yongyi Mao , Hongyu Guo , Jie Xu

Crowdsourcing is a common approach to rapidly annotate large volumes of data in machine learning applications. Typically, crowd workers are compensated with a flat rate based on an estimated completion time to meet a target hourly wage.…

人机交互 · 计算机科学 2024-12-03 Gordon Lim , Stefan Larson , Yu Huang , Kevin Leach

Labeling is onerous for crowd counting as it should annotate each individual in crowd images. Recently, several methods have been proposed for semi-supervised crowd counting to reduce the labeling efforts. Given a limited labeling budget,…

计算机视觉与模式识别 · 计算机科学 2021-08-09 Yongtuo Liu , Sucheng Ren , Liangyu Chai , Hanjie Wu , Jing Qin , Dan Xu , Shengfeng He

Due to concerns about human error in crowdsourcing, it is standard practice to collect labels for the same data point from multiple internet workers. We here show that the resulting budget can be used more effectively with a flexible worker…

人机交互 · 计算机科学 2019-01-29 Mehrnoosh Sameki , Sha Lai , Kate K. Mays , Lei Guo , Prakash Ishwar , Margrit Betke

In representation learning, there has been recent interest in developing algorithms to disentangle the ground-truth generative factors behind a dataset, and metrics to quantify how fully this occurs. However, these algorithms and metrics…

机器学习 · 计算机科学 2022-04-11 Andrew Slavin Ross , Finale Doshi-Velez