中文
相关论文

相关论文: Streaming Bayesian Inference for Crowdsourced Clas…

200 篇论文

Crowdwork often entails tackling cognitively-demanding and time-consuming tasks. Crowdsourcing can be used for complex annotation tasks, from medical imaging to geospatial data, and such data powers sensitive applications, such as health…

人机交互 · 计算机科学 2020-09-07 Akira Matsui , Emilio Ferrara , Fred Morstatter , Andres Abeliuk , Aram Galstyan

Existing works for truth discovery in categorical data usually assume that claimed values are mutually exclusive and only one among them is correct. However, many claimed values are not mutually exclusive even for functional predicates due…

数据库 · 计算机科学 2019-04-24 Woohwan Jung , Younghoon Kim , Kyuseok Shim

The evaluation of noisy binary classifiers on unlabeled data is treated as a streaming task: given a data sketch of the decisions by an ensemble, estimate the true prevalence of the labels as well as each classifier's accuracy on them. Two…

机器学习 · 统计学 2023-09-11 Andrés Corrada-Emmanuel

We target the problem of accuracy and robustness in causal inference from finite data sets. Some state-of-the-art algorithms produce clear output complete with solid theoretical guarantees but are susceptible to propagating erroneous…

人工智能 · 计算机科学 2012-10-19 Tom Claassen , Tom Heskes

In applied statistics and machine learning, the "gold standards" used for training are often biased and almost always noisy. Dawid and Skene's justifiably popular crowdsourcing model adjusts for rater (coder, annotator) sensitivity and…

机器学习 · 计算机科学 2024-10-23 Seong Woo Han , Ozan Adıgüzel , Bob Carpenter

Sequential probabilistic inference from streaming observations requires modeling distributions over future trajectories as new observations arrive. Although diffusion and flow-matching models are effective at capturing high-dimensional,…

机器学习 · 计算机科学 2026-05-15 Yinan Huang , Hans Hao-Hsun Hsu , Junran Wang , Bo Dai , Pan Li

The spread of online misinformation poses serious threats to democratic societies. Traditionally, expert fact-checkers verify the truthfulness of information through investigative processes. However, the volume and immediacy of online…

信息检索 · 计算机科学 2025-06-12 Michael Soprano

Hierarchies of concepts are useful in many applications from navigation to organization of objects. Usually, a hierarchy is created in a centralized manner by employing a group of domain experts, a time-consuming and expensive process. The…

人工智能 · 计算机科学 2015-08-04 Yuyin Sun , Adish Singla , Dieter Fox , Andreas Krause

We study the problem of clustering a set of items from binary user feedback. Such a problem arises in crowdsourcing platforms solving large-scale labeling tasks with minimal effort put on the users. For example, in some of the recent…

机器学习 · 统计学 2024-12-20 Kaito Ariu , Jungseul Ok , Alexandre Proutiere , Se-Young Yun

Crowdsourcing offers a practical method for ranking and scoring large amounts of items. To investigate the algorithms and incentives that can be used in crowdsourcing quality evaluations, we built CrowdGrader, a tool that lets students…

社会与信息网络 · 计算机科学 2013-08-27 Luca de Alfaro , Michael Shavlovsky

Federated learning brings potential benefits of faster learning, better solutions, and a greater propensity to transfer when heterogeneous data from different parties increases diversity. However, because federated learning tasks tend to be…

机器学习 · 计算机科学 2021-01-18 Duc Thien Nguyen , Shiau Hoong Lim , Laura Wynter , Desmond Cai

Harnessing human computation for solving complex problems call spawns the issue of finding the unknown competitive group of solvers. In this paper, we propose an approach called Friendlysourcing to build up teams from social network…

社会与信息网络 · 计算机科学 2013-05-30 Iheb Ben Amor , Athman Bougetteya , Mourad Ouziri , Salima Benbernou , Mohamed Nadif

Community detection, which focuses on recovering the group structure within networks, is a crucial and fundamental task in network analysis. However, the detection process can be quite challenging and unstable when community signals are…

统计理论 · 数学 2025-03-11 Tianchen Gao , Jingyuan Liu , Rui Pan , Ao Sun

We propose a scalable Bayesian preference learning method for jointly predicting the preferences of individuals as well as the consensus of a crowd from pairwise labels. Peoples' opinions often differ greatly, making it difficult to predict…

机器学习 · 计算机科学 2019-12-13 Edwin Simpson , Iryna Gurevych

The data distribution in popular crowd counting datasets is typically heavy tailed and discontinuous. This skew affects all stages within the pipelines of deep crowd counting approaches. Specifically, the approaches exhibit unacceptably…

计算机视觉与模式识别 · 计算机科学 2022-04-12 Sravya Vardhani Shivapuja , Ashwin Gopinath , Ayush Gupta , Ganesh Ramakrishnan , Ravi Kiran Sarvadevabhatla

We introduce a novel noisy sorting model motivated by the Just Noticeable Difference (JND) model from experimental psychology. The goal of our model is to capture the low quality of the data that are collected from crowdsourcing…

数据结构与算法 · 计算机科学 2023-10-24 Ellen Vitercik , Manolis Zampetakis , David Zhang

Distant supervision is a popular method for performing relation extraction from text that is known to produce noisy labels. Most progress in relation extraction and classification has been made with crowdsourced corrections to…

计算与语言 · 计算机科学 2022-09-21 Anca Dumitrache , Lora Aroyo , Chris Welty

Discovering statistically significant patterns from databases is an important challenging problem. The main obstacle of this problem is in the difficulty of taking into account the selection bias, i.e., the bias arising from the fact that…

机器学习 · 统计学 2016-03-10 Shinya Suzumura , Kazuya Nakagawa , Mahito Sugiyama , Koji Tsuda , Ichiro Takeuchi

Spatial Crowdsourcing (SC) is a novel platform that engages individuals in the act of collecting various types of spatial data. This method of data collection can significantly reduce cost and turnover time, and is particularly useful in…

数据库 · 计算机科学 2017-04-27 Luan Tran , Hien To , Liyue Fan , Cyrus Shahabi

Crowdsourcing provides a popular paradigm for data collection at scale. We study the problem of selecting subsets of workers from a given worker pool to maximize the accuracy under a budget constraint. One natural question is whether we…

机器学习 · 统计学 2015-02-04 Hongwei Li , Qiang Liu