中文
相关论文

相关论文: Microtask crowdsourcing for disease mention annota…

200 篇论文

Social media, especially Twitter, is being increasingly used for research with predictive analytics. In social media studies, natural language processing (NLP) techniques are used in conjunction with expert-based, manual and qualitative…

计算与语言 · 计算机科学 2020-04-03 Yunpeng Zhao , Mattia Prosperi , Tianchen Lyu , Yi Guo , Jiang Bian

Recent advancements in medical imaging and artificial intelligence (AI) have greatly enhanced diagnostic capabilities, but the development of effective deep learning (DL) models is still constrained by the lack of high-quality annotated…

图像与视频处理 · 电气工程与系统科学 2025-07-22 Amir Syahmi , Xiangrong Lu , Yinxuan Li , Haoxuan Yao , Hanjun Jiang , Ishita Acharya , Shiyi Wang , Yang Nan , Xiaodan Xing , Guang Yang

Annotation quality and quantity positively affect the learning performance of sequence labeling, a vital task in Natural Language Processing. Hiring domain experts to annotate a corpus is very costly in terms of money and time.…

人机交互 · 计算机科学 2023-07-04 Nasim Sabetpour , Adithya Kulkarni , Sihong Xie , Qi Li

Large language models (LLMs) are remarkable data annotators. They can be used to generate high-fidelity supervised training data, as well as survey and experimental data. With the widespread adoption of LLMs, human gold--standard…

计算与语言 · 计算机科学 2023-06-14 Veniamin Veselovsky , Manoel Horta Ribeiro , Robert West

High throughput extraction and structured labeling of data from academic articles is critical to enable downstream machine learning applications and secondary analyses. We have embedded multimodal data curation into the academic publishing…

计算与语言 · 计算机科学 2024-09-26 Jorge Abreu-Vicente , Hannah Sonntag , Thomas Eidens , Cassie S. Mitchell , Thomas Lemberger

We present crowdsourcing as an additional modality to aid radiologists in the diagnosis of lung cancer from clinical chest computed tomography (CT) scans. More specifically, a complete workflow is introduced which can help maximize the…

计算机视觉与模式识别 · 计算机科学 2018-09-19 Saeed Boorboor , Saad Nadeem , Ji Hwan Park , Kevin Baker , Arie Kaufman

Annotated images are required for both supervised model training and evaluation in image classification. Manually annotating images is arduous and expensive, especially for multi-labeled images. A recent trend for conducting such laboursome…

计算机视觉与模式识别 · 计算机科学 2022-12-07 Jianzhe Lin , Tianze Yu , Z. Jane Wang

Recent studies indicated GPT-4 outperforms online crowd workers in data labeling accuracy, notably workers from Amazon Mechanical Turk (MTurk). However, these studies were criticized for deviating from standard crowdsourcing practices and…

Semantic relation extraction is one of the frontiers of biomedical natural language processing research. Gold standards are key tools for advancing this research. It is challenging to generate these standards because of the high cost of…

计算与语言 · 计算机科学 2015-05-26 Tong Shu Li , Benjamin M. Good , Andrew I. Su

Biomedical image analysis algorithm validation depends on high-quality annotation of reference datasets, for which labeling instructions are key. Despite their importance, their optimization remains largely unexplored. Here, we present the…

Cognitive computing systems require human labeled data for evaluation, and often for training. The standard practice used in gathering this data minimizes disagreement between annotators, and we have found this results in data that fails to…

计算与语言 · 计算机科学 2018-09-27 Anca Dumitrache , Lora Aroyo , Chris Welty

Bioinformatics workflows are essential for complex biological data analyses and are often described in scientific articles with source code in public repositories. Extracting detailed workflow information from articles can improve…

计算与语言 · 计算机科学 2025-03-11 Clémence Sebe , Sarah Cohen-Boulakia , Olivier Ferret , Aurélie Névéol

We introduce Biomed-Enriched, a biomedical text dataset constructed from PubMed via a two-stage annotation process. In the first stage, a large language model annotates 400K paragraphs from PubMed scientific articles, assigning scores for…

计算与语言 · 计算机科学 2025-06-26 Rian Touchent , Nathan Godey , Eric de la Clergerie

This paper introduces CODA-19, a human-annotated dataset that codes the Background, Purpose, Method, Finding/Contribution, and Other sections of 10,966 English abstracts in the COVID-19 Open Research Dataset. CODA-19 was created by 248…

计算与语言 · 计算机科学 2020-09-21 Ting-Hao 'Kenneth' Huang , Chieh-Yang Huang , Chien-Kuang Cornelia Ding , Yen-Chia Hsu , C. Lee Giles

Whether Large Language Models (LLMs) can outperform crowdsourcing on the data annotation task is attracting interest recently. Some works verified this issue with the average performance of individual crowd workers and LLM workers on some…

计算与语言 · 计算机科学 2024-01-19 Jiyi Li

Crowdsourcing has been the prevalent paradigm for creating natural language understanding datasets in recent years. A common crowdsourcing practice is to recruit a small number of high-quality workers, and have them massively generate…

计算与语言 · 计算机科学 2019-08-29 Mor Geva , Yoav Goldberg , Jonathan Berant

Manually annotated data is key to developing text-mining and information-extraction algorithms. However, human annotation requires considerable time, effort and expertise. Given the rapid growth of biomedical literature, it is paramount to…

人机交互 · 计算机科学 2020-04-27 Rezarta Islamaj , Dongseop Kwon , Sun Kim , Zhiyong Lu

Many NLP applications require manual data annotations for a variety of tasks, notably to train classifiers or evaluate the performance of unsupervised models. Depending on the size and degree of complexity, the tasks may be conducted by…

计算与语言 · 计算机科学 2023-07-20 Fabrizio Gilardi , Meysam Alizadeh , Maël Kubli

Crowdsourcing is a relatively economic and efficient solution to collect annotations from the crowd through online platforms. Answers collected from workers with different expertise may be noisy and unreliable, and the quality of annotated…

机器学习 · 计算机科学 2020-01-08 Jingzheng Tu , Guoxian Yu , Jun Wang , Carlotta Domeniconi , Xiangliang Zhang

To efficiently establish training databases for machine learning methods, collaborative and crowdsourcing platforms have been investigated to collectively tackle the annotation effort. However, when this concept is ported to the medical…

计算机视觉与模式识别 · 计算机科学 2017-08-22 Martin Rajchl , Lisa M. Koch , Christian Ledig , Jonathan Passerat-Palmbach , Kazunari Misawa , Kensaku Mori , Daniel Rueckert
‹ 上一页 1 2 3 10 下一页 ›