English
Related papers

Related papers: Crowdsourcing with Difficulty: A Bayesian Rating M…

200 papers

With the increased interest in machine learning and big data problems, the need for large amounts of labelled data has also grown. However, it is often infeasible to get experts to label all of this data, which leads many practitioners to…

Machine Learning · Computer Science 2021-05-31 Pierce Burke , Richard Klein

The Dawid-Skene model is the most widely assumed model in the analysis of crowdsourcing algorithms that estimate ground-truth labels from noisy worker responses. In this work, we are motivated by crowdsourcing applications where workers…

Machine Learning · Computer Science 2024-08-13 Saptarshi Mandal , Seo Taek Kong , Dimitrios Katselis , R. Srikant

Mislabeled, duplicated, or biased data in real-world scenarios can lead to prolonged training and even hinder model convergence. Traditional solutions prioritizing easy or hard samples lack the flexibility to handle such a variety…

Machine Learning · Computer Science 2023-11-08 Zhijie Deng , Peng Cui , Jun Zhu

Crowdsourcing systems have been used to accumulate massive amounts of labeled data for applications such as computer vision and natural language processing. However, because crowdsourced labeling is inherently dynamic and uncertain,…

Machine Learning · Computer Science 2023-10-26 Mohammad S. Majdi , Jeffrey J. Rodriguez

We consider the problem of aggregating predictions or measurements from a set of human forecasters, models, sensors or other instruments which may be subject to bias or miscalibration and random heteroscedastic noise. We propose a Bayesian…

Statistical Finance · Quantitative Finance 2021-01-12 Chirag Nagpal , Robert E. Tillman , Prashant Reddy , Manuela Veloso

Algorithmic recommendation based on noisy preference measurement is prevalent in recommendation systems. This paper discusses the consequences of such recommendation on market concentration and inequality. Binary types denoting a…

Theoretical Economics · Economics 2025-10-21 Andreas Haupt

Modern, state-of-the-art deep learning approaches yield human like performance in numerous object detection and classification tasks. The foundation for their success is the availability of training datasets of substantially high quantity,…

Multiple measures, such as WEAT or MAC, attempt to quantify the magnitude of bias present in word embeddings in terms of a single-number metric. However, such metrics and the related statistical significance calculations rely on treating…

Computation and Language · Computer Science 2023-06-16 Alicja Dobrzeniecka , Rafal Urbaniak

This paper presents a generic Bayesian framework that enables any deep learning model to actively learn from targeted crowds. Our framework inherits from recent advances in Bayesian deep learning, and extends existing work by considering…

Machine Learning · Computer Science 2018-03-13 Jie Yang , Thomas Drake , Andreas Damianou , Yoelle Maarek

This paper introduces mixsemble, an ensemble method that adapts the Dawid-Skene model to aggregate predictions from multiple model-based clustering algorithms. Unlike traditional crowdsourcing, which relies on human labels, the framework…

Machine Learning · Computer Science 2025-10-01 Jordyn E. A. Lorentz , Katharine M. Clark

We study the problem of clustering a set of items from binary user feedback. Such a problem arises in crowdsourcing platforms solving large-scale labeling tasks with minimal effort put on the users. For example, in some of the recent…

Machine Learning · Statistics 2024-12-20 Kaito Ariu , Jungseul Ok , Alexandre Proutiere , Se-Young Yun

We propose a new probabilistic graphical model that jointly models the difficulties of questions, the abilities of participants and the correct answers to questions in aptitude testing and crowdsourcing settings. We devise an active…

Machine Learning · Computer Science 2012-07-03 Yoram Bachrach , Thore Graepel , Tom Minka , John Guiver

Acquiring fine-grained object detection annotations in unconstrained images is time-consuming, expensive, and prone to noise, especially in crowdsourcing scenarios. Most prior object detection methods assume accurate annotations; A few…

Computer Vision and Pattern Recognition · Computer Science 2024-07-12 Zhi Qin Tan , Olga Isupova , Gustavo Carneiro , Xiatian Zhu , Yunpeng Li

Datasets for training crowd counting deep networks are typically heavy-tailed in count distribution and exhibit discontinuities across the count range. As a result, the de facto statistical measures (MSE, MAE) exhibit large variance and…

Computer Vision and Pattern Recognition · Computer Science 2021-08-20 Sravya Vardhani Shivapuja , Mansi Pradeep Khamkar , Divij Bajaj , Ganesh Ramakrishnan , Ravi Kiran Sarvadevabhatla

Crowdsourcing platforms provide marketplaces where task requesters can pay to get labels on their data. Such markets have emerged recently as popular venues for collecting annotations that are crucial in training machine learning models in…

Machine Learning · Computer Science 2017-08-28 Ashish Khetan , Sewoong Oh

Modern machine learning approaches have led to performant diagnostic models for a variety of health conditions. Several machine learning approaches, such as decision trees and deep neural networks, can, in principle, approximate any…

Human-Computer Interaction · Computer Science 2024-06-05 Peter Washington

The data distribution in popular crowd counting datasets is typically heavy tailed and discontinuous. This skew affects all stages within the pipelines of deep crowd counting approaches. Specifically, the approaches exhibit unacceptably…

Computer Vision and Pattern Recognition · Computer Science 2022-04-12 Sravya Vardhani Shivapuja , Ashwin Gopinath , Ayush Gupta , Ganesh Ramakrishnan , Ravi Kiran Sarvadevabhatla

Regression Discontinuity Design (RDD) is a popular framework for estimating a causal effect in settings where treatment is assigned if an observed covariate exceeds a fixed threshold. We consider estimation and inference in the common…

Statistics Theory · Mathematics 2025-04-16 Kevin Tao , Y. Samuel Wang , David Ruppert

Crowdsourcing provides a practical way to obtain large amounts of labeled data at a low cost. However, the annotation quality of annotators varies considerably, which imposes new challenges in learning a high-quality model from the…

Machine Learning · Computer Science 2021-06-15 Zhendong Chu , Jing Ma , Hongning Wang

In crowdsourced preference aggregation, it is often assumed that all the annotators are subject to a common preference or utility function which generates their comparison behaviors in experiments. However, in reality annotators are subject…

Human-Computer Interaction · Computer Science 2016-07-13 Qianqian Xu , Jiechao Xiong , Xiaochun Cao , Yuan Yao