中文
相关论文

相关论文: NUBES: A Corpus of Negation and Uncertainty in Spa…

200 篇论文

Disentangling conversations mixed together in a single stream of messages is a difficult task, made harder by the lack of large manually annotated datasets. We created a new dataset of 77,563 messages manually annotated with reply-structure…

Online medical literature has made health information more available than ever, however, the barrier of complex medical jargon prevents the general public from understanding it. Though parallel and comparable corpora for Biomedical Text…

计算与语言 · 计算机科学 2025-06-17 William Xia , Ishita Unde , Brian Ondov , Dina Demner-Fushman

Existing datasets for automated fact-checking have substantial limitations, such as relying on artificial claims, lacking annotations for evidence and intermediate reasoning, or including evidence published after the claim. In this paper we…

计算与语言 · 计算机科学 2023-11-09 Michael Schlichtkrull , Zhijiang Guo , Andreas Vlachos

We computed both Word and Sub-word Embeddings using FastText. For Sub-word embeddings we selected Byte Pair Encoding (BPE) algorithm to represent the sub-words. We evaluated the Biomedical Word Embeddings obtaining better results than…

This paper summarizes the main findings of the ADoBo 2021 shared task, proposed in the context of IberLef 2021. In this task, we invited participants to detect lexical borrowings (coming mostly from English) in Spanish newswire texts. This…

The quality of training data is one of the crucial problems when a learning-centered approach is employed. This paper proposes a new method to investigate the quality of a large corpus designed for the recognizing textual entailment (RTE)…

计算与语言 · 计算机科学 2018-04-24 Masatoshi Tsuchiya

Code-switching (CS) remains a significant challenge in Natural Language Processing (NLP), mainly due a lack of relevant data. In the context of the contact between the Basque and Spanish languages in the north of the Iberian Peninsula, CS…

计算与语言 · 计算机科学 2025-02-06 Maite Heredia , Jeremy Barnes , Aitor Soroa

Move structures have been studied in English for Specific Purposes (ESP) and English for Academic Purposes (EAP) for decades. However, there are few move annotation corpora for Research Article (RA) abstracts. In this paper, we introduce…

计算与语言 · 计算机科学 2024-03-26 Hongzheng Li , Ruojin Wang , Ge Shi , Xing Lv , Lei Lei , Chong Feng , Fang Liu , Jinkun Lin , Yangguang Mei , Lingnan Xu

In this paper, we reflect on ways to improve the quality of bio-medical information retrieval by drawing implicit negative feedback from negated information in noisy natural language search queries. We begin by studying the extent to which…

信息检索 · 计算机科学 2016-08-08 Lorenz Kuhn , Carsten Eickhoff

The Abstract Meaning Representation (AMR) formalism, designed originally for English, has been adapted to a number of languages. We build on previous work proposing the annotation of AMR in Spanish, which resulted in the release of 50…

计算与语言 · 计算机科学 2022-04-19 Shira Wein , Lucia Donatelli , Ethan Ricker , Calvin Engstrom , Alex Nelson , Nathan Schneider

Negative medical findings are prevalent in clinical reports, yet discriminating them from positive findings remains a challenging task for information extraction. Most of the existing systems treat this task as a pipeline of two separate…

计算与语言 · 计算机科学 2020-01-23 Parminder Bhatia , Busra Celikkaya , Mohammed Khalilia

Blogs are a source of grey literature which are widely adopted by software practitioners for disseminating opinion and experience. Analysing such articles can provide useful insights into the state-of-practice for software engineering…

软件工程 · 计算机科学 2021-06-22 Ashley Williams , Matthew Shardlow , Austen Rainer

Although pre-trained named entity recognition (NER) models are highly accurate on modern corpora, they underperform on historical texts due to differences in language OCR errors. In this work, we develop a new NER corpus of 3.6M sentences…

计算与语言 · 计算机科学 2023-06-08 Vít Novotný , Kristýna Luger , Michal Štefánik , Tereza Vrabcová , Aleš Horák

This paper details LTG-Oslo team's participation in the sentiment track of the NEGES 2019 evaluation campaign. We participated in the task with a hierarchical multi-task network, which used shared lower-layers in a deep BiLSTM to predict…

计算与语言 · 计算机科学 2019-06-19 Jeremy Barnes

Nowadays, with the booming development of the Internet, people benefit from its convenience due to its open and sharing nature. A large volume of natural language texts is being generated by users in various forms, such as search queries,…

计算与语言 · 计算机科学 2019-08-07 Chenwei Zhang

Many European languages possess rich biblical translation histories, yet existing corpora - in prioritizing linguistic breadth - often fail to capture this depth. To address this gap, we introduce a multilingual corpus of 651 New Testament…

计算与语言 · 计算机科学 2026-05-14 Maciej Rapacz , Aleksander Smywiński-Pohl

Over the course of the COVID-19 pandemic, large volumes of biomedical information concerning this new disease have been published on social media. Some of this information can pose a real danger to people's health, particularly when false…

计算与语言 · 计算机科学 2022-04-27 Isabelle Mohr , Amelie Wührl , Roman Klinger

Understanding causal narratives communicated in clinical notes can help make strides towards personalized healthcare. Extracted causal information from clinical notes can be combined with structured EHR data such as patients' demographics,…

计算与语言 · 计算机科学 2022-03-15 Vivek Khetan , Md Imbesat Hassan Rizvi , Jessica Huber , Paige Bartusiak , Bogdan Sacaleanu , Andrew Fano

Detecting non-factual content is a longstanding goal to increase the trustworthiness of large language models (LLMs) generations. Current factuality probes, trained using humanannotated labels, exhibit limited transferability to…

计算与语言 · 计算机科学 2024-04-11 Xiaokang Zhang , Zijun Yao , Jing Zhang , Kaifeng Yun , Jifan Yu , Juanzi Li , Jie Tang