中文
相关论文

相关论文: NUBES: A Corpus of Negation and Uncertainty in Spa…

200 篇论文

We present a novel benchmark and associated evaluation metrics for assessing the performance of text anonymization methods. Text anonymization, defined as the task of editing a text document to prevent the disclosure of personal…

计算与语言 · 计算机科学 2022-07-04 Ildikó Pilán , Pierre Lison , Lilja Øvrelid , Anthi Papadopoulou , David Sánchez , Montserrat Batet

This survey focuses in encoder Language Models for solving tasks in the clinical domain in the Spanish language. We review the contributions of 17 corpora focused mainly in clinical tasks, then list the most relevant Spanish Language Models…

计算与语言 · 计算机科学 2023-08-07 Guillem García Subies , Álvaro Barbero Jiménez , Paloma Martínez Fernández

Scarcity of resources such as annotated text corpora for under-resourced languages like Albanian is a serious impediment in computational linguistics and natural language processing research. This paper presents AlbNER, a corpus of 900…

计算与语言 · 计算机科学 2023-09-19 Erion Çano

Extracting drug use information from unstructured Electronic Health Records remains a major challenge in clinical Natural Language Processing. While Large Language Models demonstrate advancements, their use in clinical NLP is limited by…

计算与语言 · 计算机科学 2026-02-11 Huu-Huy-Hoang Tran , Gia-Bao Duong , Quoc-Viet-Anh Tran , Thi-Hai-Yen Vuong , Hoang-Quynh Le

As large language models are increasingly deployed for clinical text, ensuring they can reliably signal their own uncertainty becomes critical. Most existing uncertainty quantification (UQ) methods are designed for open-domain generation…

计算与语言 · 计算机科学 2026-05-28 Bushi Xiao , Sarvesh Soni , Daisy Zhe Wang

Objective: To build a comprehensive corpus covering syntactic and semantic annotations of Chinese clinical texts with corresponding annotation guidelines and methods as well as to develop tools trained on the annotated corpus, which…

计算与语言 · 计算机科学 2016-11-09 Bin He , Bin Dong , Yi Guan , Jinfeng Yang , Zhipeng Jiang , Qiubin Yu , Jianyi Cheng , Chunyan Qu

When speaking or writing, people omit information that seems clear and evident, such that only part of the message is expressed in words. Especially in argumentative texts it is very common that (important) parts of the argument are implied…

计算与语言 · 计算机科学 2019-12-24 Maria Becker , Katharina Korfhage , Anette Frank

This paper analyzes negation in eight popular corpora spanning six natural language understanding tasks. We show that these corpora have few negations compared to general-purpose English, and that the few negations in them are often…

计算与语言 · 计算机科学 2022-03-18 Md Mosharaf Hossain , Dhivya Chinnappa , Eduardo Blanco

We present NoticIA, a dataset consisting of 850 Spanish news articles featuring prominent clickbait headlines, each paired with high-quality, single-sentence generative summarizations written by humans. This task demands advanced text…

计算与语言 · 计算机科学 2024-06-03 Iker García-Ferrero , Begoña Altuna

We present Stanceosaurus, a new corpus of 28,033 tweets in English, Hindi, and Arabic annotated with stance towards 251 misinformation claims. As far as we are aware, it is the largest corpus annotated with stance towards misinformation…

计算与语言 · 计算机科学 2022-10-31 Jonathan Zheng , Ashutosh Baheti , Tarek Naous , Wei Xu , Alan Ritter

Medical image segmentation modeling is a high-stakes task where understanding of uncertainty is crucial for addressing visual ambiguity. Prior work has developed segmentation models utilizing probabilistic or generative mechanisms to infer…

计算机视觉与模式识别 · 计算机科学 2025-06-13 Andre Ye , Quan Ze Chen , Amy Zhang

The performance and usability of Large-Language Models (LLMs) are driving their use in explanation generation tasks. However, despite their widespread adoption, LLM explanations have been found to be unreliable, making it difficult for…

We construct the first ever multimodal sarcasm dataset for Spanish. The audiovisual dataset consists of sarcasm annotated text that is aligned with video and audio. The dataset represents two varieties of Spanish, a Latin American variety…

计算与语言 · 计算机科学 2021-05-13 Khalid Alnajjar , Mika Hämäläinen

We present an evaluation of text simplification (TS) in Spanish for a production system, by means of two corpora focused in both complex-sentence and complex-word identification. We compare the most prevalent Spanish-specific readability…

计算与语言 · 计算机科学 2023-08-16 Adrian de Wynter , Anthony Hevia , Si-Qing Chen

This paper describes a novel protocol for collecting speech data from subjects induced to have different degrees of trust in the skills of a conversational agent. The protocol consists of an interactive session where the subject is asked to…

Word embeddings are a popular approach to unsupervised learning of word relationships that are widely used in natural language processing. In this article, we present a new set of embeddings for medical concepts learned using an extremely…

This paper addresses the problem of fake news detection in Spanish using Machine Learning techniques. It is fundamentally the same problem tackled for the English language; however, there is not a significant amount of publicly available…

计算与语言 · 计算机科学 2021-10-14 Kevin Martínez-Gallego , Andrés M. Álvarez-Ortiz , Julián D. Arias-Londoño

Sentence embeddings have become an essential part of today's natural language processing (NLP) systems, especially together advanced deep learning methods. Although pre-trained sentence encoders are available in the general domain, none…

计算与语言 · 计算机科学 2020-01-28 Qingyu Chen , Yifan Peng , Zhiyong Lu

We develop novel annotation guidelines for sentence-level subjectivity detection, which are not limited to language-specific cues. We use our guidelines to collect NewsSD-ENG, a corpus of 638 objective and 411 subjective sentences extracted…

The RareDis corpus contains more than 5,000 rare diseases and almost 6,000 clinical manifestations are annotated. Moreover, the Inter Annotator Agreement evaluation shows a relatively high agreement (F1-measure equal to 83.5% under exact…