中文
相关论文

相关论文: Exact Exponent in Optimal Rates for Crowdsourcing

200 篇论文

There is a rapidly increasing interest in crowdsourcing for data labeling. By crowdsourcing, a large number of labels can be often quickly gathered at low cost. However, the labels provided by the crowdsourcing workers are usually not of…

机器学习 · 计算机科学 2015-03-26 Dengyong Zhou , Qiang Liu , John C. Platt , Christopher Meek , Nihar B. Shah

This paper introduces mixsemble, an ensemble method that adapts the Dawid-Skene model to aggregate predictions from multiple model-based clustering algorithms. Unlike traditional crowdsourcing, which relies on human labels, the framework…

机器学习 · 计算机科学 2025-10-01 Jordyn E. A. Lorentz , Katharine M. Clark

We study crowdsourcing quality management, that is, given worker responses to a set of tasks, our goal is to jointly estimate the true answers for the tasks, as well as the quality of the workers. Prior work on this problem relies primarily…

其他计算机科学 · 计算机科学 2015-03-03 Akash Das Sarma , Aditya Parameswaran , Jennifer Widom

We consider ensembles of channel codes that are partitioned into bins, and focus on analysis of exact random coding error exponents associated with optimum decoding of the index of the bin to which the transmitted codeword belongs. Two main…

信息论 · 计算机科学 2016-11-18 Neri Merhav

We consider worker skill estimation for the single-coin Dawid-Skene crowdsourcing model. In practice, skill-estimation is challenging because worker assignments are sparse and irregular due to the arbitrary and uncontrolled availability of…

人机交互 · 计算机科学 2020-07-24 Yao Ma , Alex Olshevsky , Venkatesh Saligrama , Csaba Szepesvari

Crowdsourcing platforms emerged as popular venues for purchasing human intelligence at low cost for large volume of tasks. As many low-paid workers are prone to give noisy answers, a common practice is to add redundancy by assigning…

机器学习 · 计算机科学 2018-10-09 Jungseul Ok , Sewoong Oh , Yunhun Jang , Jinwoo Shin , Yung Yi

This paper considers the problem of channel coding with a given (possibly suboptimal) maximum-metric decoding rule. A cost-constrained random-coding ensemble with multiple auxiliary costs is introduced, and is shown to achieve error…

信息论 · 计算机科学 2014-03-05 Jonathan Scarlett , Alfonso Martinez , Albert Guillén i Fàbregas

The minimum error entropy (MEE) criterion has been successfully used in fields such as parameter estimation, system identification and the supervised machine learning. There is in general no explicit expression for the optimal MEE estimate…

信息论 · 计算机科学 2015-04-14 Badong Chen , Guangmin Wang , Nanning Zheng , Jose C. Principe

We consider the problem of ranking $n$ experts according to their abilities, based on the correctness of their answers to $d$ questions. This is modeled by the so-called crowd-sourcing model, where the answer of expert $i$ on question $k$…

统计理论 · 数学 2025-12-25 Alexandra Carpentier , Nicolas Verzelen

Crowdsourcing provides a popular paradigm for data collection at scale. We study the problem of selecting subsets of workers from a given worker pool to maximize the accuracy under a budget constraint. One natural question is whether we…

机器学习 · 统计学 2015-02-04 Hongwei Li , Qiang Liu

Crowdsourcing platforms enable to propose simple human intelligence tasks to a large number of participants who realise these tasks. The workers often receive a small amount of money or the platforms include some other incentive mechanisms,…

人工智能 · 计算机科学 2016-10-03 Amal Ben Rjab , Mouloud Kharoune , Zoltan Miklos , Arnaud Martin

We study a problem of optimal information gathering from multiple data providers that need to be incentivized to provide accurate information. This problem arises in many real world applications that rely on crowdsourced data sets, but…

计算机科学与博弈论 · 计算机科学 2017-11-27 Goran Radanovic , Adish Singla , Andreas Krause , Boi Faltings

We propose a streaming algorithm for the binary classification of data based on crowdsourcing. The algorithm learns the competence of each labeller by comparing her labels to those of other labellers on the same tasks and uses this…

机器学习 · 统计学 2016-02-24 Thomas Bonald , Richard Combes

Crowdsourcing is a relatively economic and efficient solution to collect annotations from the crowd through online platforms. Answers collected from workers with different expertise may be noisy and unreliable, and the quality of annotated…

机器学习 · 计算机科学 2020-01-08 Jingzheng Tu , Guoxian Yu , Jun Wang , Carlotta Domeniconi , Xiangliang Zhang

We study the expectation-maximization (EM) algorithm for general latent-variable models under (i) distributional misspecification and (ii) nonidentifiability induced by a group action. We formulate EM on the quotient parameter space and…

统计理论 · 数学 2026-01-06 Koustav Mallik

In mixtures-of-experts (ME) model, where a number of submodels (experts) are combined, there have been two longstanding problems: (i) how many experts should be chosen, given the size of the training data? (ii) given the total number of…

统计理论 · 数学 2011-11-02 Eduardo F. Mendes , Wenxin Jiang

Crowdsourcing has become a popular method for collecting labeled training data. However, in many practical scenarios traditional labeling can be difficult for crowdworkers (for example, if the data is high-dimensional or unintuitive, or the…

机器学习 · 统计学 2017-12-14 Tom Hope , Dafna Shahaf

In recent years crowdsourcing has become the method of choice for gathering labeled training data for learning algorithms. Standard approaches to crowdsourcing view the process of acquiring labeled data separately from the process of…

机器学习 · 计算机科学 2017-04-17 Pranjal Awasthi , Avrim Blum , Nika Haghtalab , Yishay Mansour

Finding a small set of representatives from an unlabeled dataset is a core problem in a broad range of applications such as dataset summarization and information extraction. Classical exemplar selection methods such as $k$-medoids work…

机器学习 · 计算机科学 2020-06-09 Chong You , Chi Li , Daniel P. Robinson , Rene Vidal

Modern machine learning algorithms need large datasets to be trained. Crowdsourcing has become a popular approach to label large datasets in a shorter time as well as at a lower cost comparing to that needed for a limited number of experts.…