中文
相关论文

相关论文: NUBES: A Corpus of Negation and Uncertainty in Spa…

200 篇论文

Massive digital data processing provides a wide range of opportunities and benefits, but at the cost of endangering personal data privacy. Anonymisation consists in removing or replacing sensitive information from data, enabling its…

计算与语言 · 计算机科学 2020-03-18 Aitor García-Pablos , Naiara Perez , Montse Cuadros

In this article we present UNED-ACCESS 2024, a bilingual dataset that consists of 1003 multiple-choice questions of university entrance level exams in Spanish and English. Questions are originally formulated in Spanish and translated…

We show-case an application of information extraction methods, such as named entity recognition (NER) and relation extraction (RE) to a novel corpus, consisting of documents, issued by a state agency. The main challenges of this corpus are:…

In this paper, we present a corpus for use in automatic readability assessment and automatic text simplification of German. The corpus is compiled from web sources and consists of approximately 211,000 sentences. As a novel contribution, it…

计算与语言 · 计算机科学 2019-09-20 Alessia Battisti , Sarah Ebling

Clinical dataset labels are rarely certain as annotators disagree and confidence is not uniform across cases. Typical aggregation procedures, such as majority voting, obscure this variability. In simple experiments on medical imaging…

The Novelties corpus is a collection of novels (and parts of novels) annotated for Named Entity Recognition (NER) among other tasks. This document describes the guidelines applied during its annotation. It contains the instructions used by…

计算与语言 · 计算机科学 2024-10-07 Arthur Amalvy , Vincent Labatut

There is a growing interest in creating tools to assist in clinical note generation using the audio of provider-patient encounters. Motivated by this goal and with the help of providers and medical scribes, we developed an annotation scheme…

Stance detection aims to determine the attitude of a given text with respect to a specific topic or claim. While stance detection has been fairly well researched in the last years, most the work has been focused on English. This is mainly…

计算与语言 · 计算机科学 2020-04-02 Elena Zotova , Rodrigo Agerri , Manuel Nuñez , German Rigau

This paper presents the InScript corpus (Narrative Texts Instantiating Script structure). InScript is a corpus of 1,000 stories centered around 10 different scenarios. Verbs and noun phrases are annotated with event and participant types,…

计算与语言 · 计算机科学 2017-03-16 Ashutosh Modi , Tatjana Anikina , Simon Ostermann , Manfred Pinkal

Fact verification has attracted a lot of attention in the machine learning and natural language processing communities, as it is one of the key methods for detecting misinformation. Existing large-scale benchmarks for this task have focused…

Annotation pipelines in Natural Language Processing (NLP) commonly assume a single latent ground truth per instance and resolve disagreement through label aggregation. Perspectivist approaches challenge this view by treating disagreement as…

计算与语言 · 计算机科学 2026-04-23 Olga Kellert , Sriya Kondury , Candice Koo , Nemika Tyagi , Steffen Eikenberry

This document, based on feedback from UMR TETIS members and the scientific literature, provides a generic methodology for creating annotation guidelines and annotated textual datasets (corpora). It covers methodological aspects, as well as…

信息检索 · 计算机科学 2026-01-21 Bahdja Boudoua , Nadia Guiffant , Mathieu Roche , Maguelonne Teisseire , Annelise Tran

In this paper, we present a new dataset and benchmark tailored to the task of semantic similarity in song lyrics. Our dataset, originally consisting of 2775 pairs of Spanish songs, was annotated in a collective annotation experiment by 63…

The ability to summarize long documents succinctly is increasingly important in daily life due to information overload, yet there is a notable lack of such summaries for Spanish documents in general, and in the legal domain in particular.…

UnScientify, a system designed to detect scientific uncertainty in scholarly full text. The system utilizes a weakly supervised technique to identify verbally expressed uncertainty in scientific texts and their authorial references. The…

计算与语言 · 计算机科学 2025-04-10 Panggih Kusuma Ningrum , Philipp Mayr , Nina Smirnova , Iana Atanassova

Negative and uncertain medical findings are frequent in radiology reports, but discriminating them from positive findings remains challenging for information extraction. Here, we propose a new algorithm, NegBio, to detect negative and…

计算与语言 · 计算机科学 2017-12-29 Yifan Peng , Xiaosong Wang , Le Lu , Mohammadhadi Bagheri , Ronald Summers , Zhiyong Lu

This paper introduces the Multi-Genre Natural Language Inference (MultiNLI) corpus, a dataset designed for use in the development and evaluation of machine learning models for sentence understanding. In addition to being one of the largest…

计算与语言 · 计算机科学 2018-02-21 Adina Williams , Nikita Nangia , Samuel R. Bowman

We present a new corpus with coreference annotation, Russian Coreference Corpus (RuCoCo). The goal of RuCoCo is to obtain a large number of annotated texts while maintaining high inter-annotator agreement. RuCoCo contains news texts in…

计算与语言 · 计算机科学 2022-06-13 Vladimir Dobrovolskii , Mariia Michurina , Alexandra Ivoylova

Resolving the scope of a negation within a sentence is a challenging NLP task. The complexity of legal texts and the lack of annotated in-domain negation corpora pose challenges for state-of-the-art (SotA) models when performing negation…

计算与语言 · 计算机科学 2023-09-19 Ramona Christen , Anastassia Shaitarova , Matthias Stürmer , Joel Niklaus

Speculation is a naturally occurring phenomena in textual data, forming an integral component of many systems, especially in the biomedical information retrieval domain. Previous work addressing cue detection and scope resolution (the two…

计算与语言 · 计算机科学 2020-01-10 Benita Kathleen Britto , Aditya Khandelwal