English
Related papers

Related papers: Easy, Reproducible and Quality-Controlled Data Col…

200 papers

Fine-grained Entity Typing is a tough task which suffers from noise samples extracted from distant supervision. Thousands of manually annotated samples can achieve greater performance than millions of samples generated by the previous…

Artificial Intelligence · Computer Science 2019-06-14 Sheng Lin , Luye Zheng , Bo Chen , Siliang Tang , Yueting Zhuang , Fei Wu , Zhigang Chen , Guoping Hu , Xiang Ren

Public omics databases like the Gene Expression Omnibus and the Sequence Read Archive offer substantial opportunities for data reuse to address novel biomedical questions. However, it is still difficult to find samples and studies of…

Coreference annotation is an important, yet expensive and time consuming, task, which often involved expert annotators trained on complex decision guidelines. To enable cheaper and more efficient annotation, we present CoRefi, a web-based…

Computation and Language · Computer Science 2020-10-07 Aaron Bornstein , Arie Cattan , Ido Dagan

Despite a significant improvement in the educational aids in terms of effective teaching-learning process, most of the educational content available to the students is less than optimal in the context of being up-to-date, exhaustive and…

Computers and Society · Computer Science 2015-08-12 Anamika Chhabra , S. R. S. Iyengar , Poonam Saini , Rajesh Shreedhar Bhat

Visualizing NLP annotation is useful for the collection of training data for the statistical NLP approaches. Existing toolkits either provide limited visual aid, or introduce comprehensive operators to realize sophisticated linguistic…

Computation and Language · Computer Science 2015-08-26 Hanchuan Li , Haichen Shen , Shengliang Xu , Congle Zhang

Sample efficiency is a crucial problem in deep reinforcement learning. Recent algorithms, such as REDQ and DroQ, found a way to improve the sample efficiency by increasing the update-to-data (UTD) ratio to 20 gradient update steps on the…

Machine Learning · Computer Science 2024-03-26 Aditya Bhatt , Daniel Palenicek , Boris Belousov , Max Argus , Artemij Amiranashvili , Thomas Brox , Jan Peters

The rapid entry of machine learning approaches in our daily activities and high-stakes domains demands transparency and scrutiny of their fairness and reliability. To help gauge machine learning models' robustness, research typically…

Machine Learning · Computer Science 2023-09-28 Oana Inel , Tim Draws , Lora Aroyo

With more and more cultural heritage data being published online, their usefulness in this open context depends on the quality and diversity of descriptive metadata for collection objects. In many cases, existing metadata is not adequate…

Computers and Society · Computer Science 2017-09-28 Chris Dijkshoorn , Victor De Boer , Lora Aroyo , Guus Schreiber

Crowdsourcing provides an efficient label collection schema for supervised machine learning. However, to control annotation cost, each instance in the crowdsourced data is typically annotated by a small number of annotators. This creates a…

Machine Learning · Computer Science 2021-07-23 Zhendong Chu , Hongning Wang

Annotations allow users to associate additional information with existing resources. Using proprietary and closed systems on the Web, users are already able to annotate multimedia resources such as images, audio and video. So far, however,…

Digital Libraries · Computer Science 2011-06-28 Bernhard Haslhofer , Rainer Simon , Robert Sanderson , Herbert van de Sompel

Crowdsourced data supports real-time decision-making but faces challenges like misinformation, errors, and contributor power concentration. This study systematically examines trust management practices across platforms categorised as…

Social and Information Networks · Computer Science 2025-11-06 Iffat Gheyas , Muhammad Rizwan Asghar , Steve Schneider , Alan Woodward

Big data have the characteristics of enormous volume, high velocity, diversity, value-sparsity, and uncertainty, which lead the knowledge learning from them full of challenges. With the emergence of crowdsourcing, versatile information can…

Machine Learning · Computer Science 2022-06-22 Jing Zhang

The evolution of AI is advancing rapidly, creating both challenges and opportunities for industry-community collaboration. In this work, we present a novel methodology aiming to facilitate this collaboration through crowdsourcing of AI…

Machine Learning · Computer Science 2021-03-26 Parthasarathy Suryanarayanan , Sundar Saranathan , Shilpa Mahatma , Divya Pathak

Labelling user data is a central part of the design and evaluation of pervasive systems that aim to support the user through situation-aware reasoning. It is essential both in designing and training the system to recognise and reason about…

Quality control in crowdsourcing systems is crucial. It is typically done after data collection, often using additional crowdsourced tasks to assess and improve the quality. These post-hoc methods can easily add cost and latency to the…

Databases · Computer Science 2017-04-11 Hamed Nilforoshan , Jiannan Wang , Eugene Wu

This paper does not describe a novel method. Instead, it studies an essential foundation for reliable benchmarking and ultimately real-world application of AI-based image analysis: generating high-quality reference annotations. Previous…

Computer Vision and Pattern Recognition · Computer Science 2024-07-29 Tim Rädsch , Annika Reinke , Vivienn Weru , Minu D. Tizabi , Nicholas Heller , Fabian Isensee , Annette Kopp-Schneider , Lena Maier-Hein

Crowdsourcing is the primary means to generate training data at scale, and when combined with sophisticated machine learning algorithms, crowdsourcing is an enabler for a variety of emergent automated applications impacting all spheres of…

Human-Computer Interaction · Computer Science 2016-10-19 Aditya Parameswaran , Akash Das Sarma , Vipul Venkataraman

Many Natural Language Processing (NLP) systems use annotated corpora for training and evaluation. However, labeled data is often costly to obtain and scaling annotation projects is difficult, which is why annotation tasks are often…

Summary descriptions of subroutines are short (usually one-sentence) natural language explanations of a subroutine's behavior and purpose in a program. These summaries are ubiquitous in documentation, and many tools such as JavaDocs and…

Software Engineering · Computer Science 2019-12-24 Zachary Eberhart , Alexander LeClair , Collin McMillan

Annotation Query (AQ) is a program that provides the ability to query many different types of NLP annotations on a text, as well as the original content and structure of the text. The query results may provide new annotations, or they may…

Distributed, Parallel, and Cluster Computing · Computer Science 2018-02-05 Darin McBeath , Ron Daniel