中文
相关论文

相关论文: Towards a Shared Rubric for Dataset Annotation

200 篇论文

This paper studies the evaluation of policies that recommend an ordered set of items (e.g., a ranking) based on some context---a common scenario in web search, ads, and recommendation. We build on techniques from combinatorial bandits to…

In conference peer review, reviewers are often asked to provide "bids" on each submitted paper that express their interest in reviewing that paper. A paper assignment algorithm then uses these bids (along with other data) to compute a…

密码学与安全 · 计算机科学 2023-03-14 Steven Jecmen , Minji Yoon , Vincent Conitzer , Nihar B. Shah , Fei Fang

Crowdsourcing provides a practical way to obtain large amounts of labeled data at a low cost. However, the annotation quality of annotators varies considerably, which imposes new challenges in learning a high-quality model from the…

机器学习 · 计算机科学 2021-06-15 Zhendong Chu , Jing Ma , Hongning Wang

Motivated by applications such as voluntary carbon markets and educational testing, we consider a market for goods with varying but hidden levels of quality in the presence of a third-party certifier. The certifier can provide informative…

计算机科学与博弈论 · 计算机科学 2023-02-01 Andreas A. Haupt , Nicole Immorlica , Brendan Lucier

Creating a linguistic resource is often done by using a machine learning model that filters the content that goes through to a human annotator, before going into the final resource. However, budgets are often limited, and the amount of…

计算与语言 · 计算机科学 2018-07-19 Filip Klubička , Giancarlo D. Salton , John D. Kelleher

Annotations allow users to associate additional information with existing resources. Using proprietary and closed systems on the Web, users are already able to annotate multimedia resources such as images, audio and video. So far, however,…

数字图书馆 · 计算机科学 2011-06-28 Bernhard Haslhofer , Rainer Simon , Robert Sanderson , Herbert van de Sompel

Most of the work in the auction design literature assumes that bidders behave rationally based on the information available for every individual auction, and the revelation principle enables designers to restrict their efforts to incentive…

计算机科学与博弈论 · 计算机科学 2024-05-14 Juncheng Li , Pingzhong Tang

Human data labeling is an important and expensive task at the heart of supervised learning systems. Hierarchies help humans understand and organize concepts. We ask whether and how concept hierarchies can inform the design of annotation…

人机交互 · 计算机科学 2023-02-24 Rickard Stureborg , Bhuwan Dhingra , Jun Yang

For data pricing, data quality is a factor that must be considered. To keep the fairness of data market from the aspect of data quality, we proposed a fair data market that considers data quality while pricing. To ensure fairness, we first…

数据库 · 计算机科学 2018-08-07 Dan Zhang , Hongzhi Wang , Xiaoou Ding , Yice Zhang , Jianzhong Li , Hong Gao

Ranking plays a central role in connecting users and providers in Information Retrieval (IR) systems, making provider-side fairness an important challenge. While recent research has begun to address fairness in ranking, most existing…

信息检索 · 计算机科学 2026-02-03 Yiteng Tu , Weihang Su , Shuguang Han , Yiqun Liu , Qingyao Ai

In this paper, we address the limitations of the common data annotation and training methods for objective single-label classification tasks. Typically, when annotating such tasks annotators are only asked to provide a single label for each…

计算与语言 · 计算机科学 2023-11-10 Ben Wu , Yue Li , Yida Mu , Carolina Scarton , Kalina Bontcheva , Xingyi Song

Human-performed annotation of sentences in legal documents is an important prerequisite to many machine learning based systems supporting legal tasks. Typically, the annotation is done sequentially, sentence by sentence, which is often time…

计算与语言 · 计算机科学 2021-12-23 Hannes Westermann , Jaromir Savelka , Vern R. Walker , Kevin D. Ashley , Karim Benyekhlef

Ranking is fundamental to many areas, such as search engine optimization, human feedback for language models, as well as peer grading. Crowdsourcing, which is often used for these tasks, requires proper incentivization to ensure accurate…

计算机科学与博弈论 · 计算机科学 2024-01-26 Kiriaki Frangias , Andrew Lin , Ellen Vitercik , Manolis Zampetakis

Linguistic bias in online news and social media is widespread but difficult to measure. Yet, its identification and quantification remain difficult due to subjectivity, context dependence, and the scarcity of high-quality gold-label…

信息检索 · 计算机科学 2025-12-17 Fabian Haak , Philipp Schaer

This paper develops and implements a scalable methodology for (a) estimating the noisiness of labels produced by a typical crowdsourcing semantic annotation task, and (b) reducing the resulting error of the labeling process by as much as…

计算与语言 · 计算机科学 2020-12-09 David Q. Sun , Hadas Kotek , Christopher Klein , Mayank Gupta , William Li , Jason D. Williams

Morality plays an important role in culture, identity, and emotion. Recent advances in natural language processing have shown that it is possible to classify moral values expressed in text at scale. Morality classification relies on human…

计算与语言 · 计算机科学 2022-10-17 Negar Mokhberian , Frederic R. Hopp , Bahareh Harandizadeh , Fred Morstatter , Kristina Lerman

We study the problem of auction design for advertising platforms that face strategic advertisers who are bidding across platforms. Each advertiser's goal is to maximize their total value or conversions while satisfying some constraint(s)…

计算机科学与博弈论 · 计算机科学 2024-05-07 Gagan Aggarwal , Andres Perlroth , Ariel Schvartzman , Mingfei Zhao

Identifying the quality of free-text arguments has become an important task in the rapidly expanding field of computational argumentation. In this work, we explore the challenging task of argument quality ranking. To this end, we created a…

计算与语言 · 计算机科学 2019-11-27 Shai Gretz , Roni Friedman , Edo Cohen-Karlik , Assaf Toledo , Dan Lahav , Ranit Aharonov , Noam Slonim

Rankings are the primary interface through which many online platforms match users to items (e.g. news, products, music, video). In these two-sided markets, not only the users draw utility from the rankings, but the rankings also determine…

信息检索 · 计算机科学 2020-06-01 Marco Morik , Ashudeep Singh , Jessica Hong , Thorsten Joachims

As larger and more comprehensive datasets become standard in contemporary machine learning, it becomes increasingly more difficult to obtain reliable, trustworthy label information with which to train sophisticated models. To address this…

机器学习 · 计算机科学 2021-06-08 Glenn Dawson , Robi Polikar