中文
相关论文

相关论文: PreCog: Improving Crowdsourced Data Quality Before…

200 篇论文

Efficient data selection is crucial to accelerate the pretraining of language model (LMs). While various methods have been proposed to enhance data efficiency, limited research has addressed the inherent conflicts between these approaches…

In this work, we initiate the investigation of optimization opportunities in collaborative crowdsourcing. Many popular applications, such as collaborative document editing, sentence translation, or citizen science resort to this special…

With the increasing pervasiveness of algorithms across industry and government, a growing body of work has grappled with how to understand their societal impact and ethical implications. Various methods have been used at different stages of…

计算机与社会 · 计算机科学 2022-07-21 Julia Barnett , Nicholas Diakopoulos

A limitation of current neural dialog models is that they tend to suffer from a lack of specificity and informativeness in generated responses, primarily due to dependence on training data that covers a limited variety of scenarios and…

计算与语言 · 计算机科学 2022-03-23 Bodhisattwa Prasad Majumder , Harsh Jhamtani , Taylor Berg-Kirkpatrick , Julian McAuley

Retrieval-Augmented Generation (RAG) has emerged as a promising paradigm for improving the timeliness of knowledge updates and the factual accuracy of large language models. However, incorporating a large volume of retrieved documents…

计算与语言 · 计算机科学 2026-05-29 Ziqiang Cui , Yunpeng Weng , Xing Tang , Peiyang Liu , Shiwei Li , Bowei He , Jiamin Chen , Yansen Zhang , Xiuqiang He , Chen Ma

Quality pretraining data is often seen as the key to high-performance language models. However, progress in understanding pretraining data has been slow due to the costly pretraining runs required for data selection experiments. We present…

计算与语言 · 计算机科学 2025-03-11 Tristan Thrush , Christopher Potts , Tatsunori Hashimoto

This paper introduces a novel crowdsourcing worker selection algorithm, enhancing annotation quality and reducing costs. Unlike previous studies targeting simpler tasks, this study contends with the complexities of label interdependencies…

计算与语言 · 计算机科学 2024-07-30 Yujie Wang , Chao Huang , Liner Yang , Zhixuan Fang , Yaping Huang , Yang Liu , Jingsi Yu , Erhong Yang

Machine learning systems are increasingly deployed in high-stakes domains, yet they remain vulnerable to bias systematic disparities that disproportionately impact specific demographic groups. Traditional bias detection methods often depend…

机器学习 · 计算机科学 2025-06-16 Chirudeep Tupakula , Rittika Shamsuddin

Rank aggregation through crowdsourcing has recently gained significant attention, particularly in the context of listwise ranking annotations. However, existing methods primarily focus on a single problem and partial ranks, while the…

机器学习 · 计算机科学 2024-10-11 Wenshui Luo , Haoyu Liu , Yongliang Ding , Tao Zhou , Sheng wan , Runze Wu , Minmin Lin , Cong Zhang , Changjie Fan , Chen Gong

As the number of individuals in a crowd grows, enumeration-based techniques become increasingly infeasible and their estimates increasingly unreliable. We propose instead an estimation-based version of the problem: we label Rough Crowd…

计算机视觉与模式识别 · 计算机科学 2025-04-21 Shengqin Jiang , Linfei Li , Haokui Zhang , Qingshan Liu , Amin Beheshti , Jian Yang , Anton van den Hengel , Quan Z. Sheng , Yuankai Qi

We study crowdsourcing quality management, that is, given worker responses to a set of tasks, our goal is to jointly estimate the true answers for the tasks, as well as the quality of the workers. Prior work on this problem relies primarily…

其他计算机科学 · 计算机科学 2015-03-03 Akash Das Sarma , Aditya Parameswaran , Jennifer Widom

Very recently crowdsourcing has become the de facto platform for distributing and collecting human computation for a wide range of tasks and applications such as information retrieval, natural language processing and machine learning.…

机器学习 · 计算机科学 2013-05-21 Ittai Abraham , Omar Alonso , Vasilis Kandylas , Aleksandrs Slivkins

Existing sequential recommendation methods rely on large amounts of training data and usually suffer from the data sparsity problem. To tackle this, the pre-training mechanism has been widely adopted, which attempts to leverage large-scale…

信息检索 · 计算机科学 2021-02-23 Chaojun Xiao , Ruobing Xie , Yuan Yao , Zhiyuan Liu , Maosong Sun , Xu Zhang , Leyu Lin

Self-supervised learning solves pretext prediction tasks that do not require annotations to learn feature representations. For vision tasks, pretext tasks such as predicting rotation, solving jigsaw are solely created from the input data.…

计算机视觉与模式识别 · 计算机科学 2021-07-01 Prashant Bhat , Elahe Arani , Bahram Zonooz

Crowd-sourcing is a cheap and popular means of creating training and evaluation datasets for machine learning, however it poses the problem of `truth inference', as individual workers cannot be wholly trusted to provide reliable…

机器学习 · 计算机科学 2019-02-26 Yuan Li , Benjamin I. P. Rubinstein , Trevor Cohn

Digitization of historical documents is a challenging task in many digital humanities projects. A popular approach for digitization is to scan the documents into images, and then convert images into text using Optical Character Recognition…

人机交互 · 计算机科学 2023-08-01 Omri Suissa , Avshalom Elmalech , Maayan Zhitomirsky-Geffet

Crowdsourcing is a multidisciplinary research area including disciplines like artificial intelligence, human-computer interaction, database, and social science. To facilitate cooperation across disciplines, reproducibility is a crucial…

数据库 · 计算机科学 2016-09-06 Ruochen Jiang , Jiannan Wang

A significant portion of the effort involved in advanced process control, process analytics, and machine learning involves acquiring and preparing data. Literature often emphasizes increasingly complex modelling techniques with incremental…

系统与控制 · 电气工程与系统科学 2023-04-07 Lim C. Siang , Shams Elnawawi , Lee D. Rippon , Daniel L. O'Connor , R. Bhushan Gopaluni

Progress in large language models is increasingly constrained by an evaluation bottleneck: benchmarks must be built and models run before iteration can begin. We investigate whether evaluation outcomes can be forecast before any experiments…

计算与语言 · 计算机科学 2026-02-05 Jungsoo Park , Ethan Mendes , Gabriel Stanovsky , Alan Ritter

Crowd sensing is a new paradigm that leverages pervasive sensor-equipped mobile devices to provide sensing services like forensic analysis, documenting public spaces, and collaboratively constructing statistical models. Extensive user…

社会与信息网络 · 计算机科学 2015-05-26 Jiajun Sun