中文
相关论文

相关论文: NUBES: A Corpus of Negation and Uncertainty in Spa…

200 篇论文

In this paper, we revisit the task of negation resolution, which includes the subtasks of cue detection (e.g. "not", "never") and scope resolution. In the context of previous shared tasks, a variety of evaluation metrics have been proposed.…

计算与语言 · 计算机科学 2021-09-22 Elizaveta Sineva , Stefan Grünewald , Annemarie Friedrich , Jonas Kuhn

While long-context large language models (LLMs) can technically summarize book-length documents (>100K tokens), the length and complexity of the documents have so far prohibited evaluations of input-dependent aspects like faithfulness. In…

计算与语言 · 计算机科学 2024-10-01 Yekyung Kim , Yapei Chang , Marzena Karpinska , Aparna Garimella , Varun Manjunatha , Kyle Lo , Tanya Goyal , Mohit Iyyer

Text articles with false claims, especially news, have recently become aggravating for the Internet users. These articles are in wide circulation and readers face difficulty discerning fact from fiction. Previous work on credibility…

计算与语言 · 计算机科学 2024-03-08 Nurendra Choudhary , Rajat Singh , Ishita Bindlish , Manish Shrivastava

A constant influx of new data poses a challenge in keeping the annotation in biological databases current. Most biological databases contain significant quantities of textual annotation, which often contains the richest source of knowledge.…

计算与语言 · 计算机科学 2013-08-22 Michael J. Bell , Matthew Collison , Phillip Lord

Information retrieval (IR) is the task of finding relevant documents in response to a user query. Although Spanish is the second most spoken native language, there are few Spanish IR datasets, which limits the development of information…

计算与语言 · 计算机科学 2025-11-20 Francisco Valentini , Viviana Cotik , Damián Furman , Ivan Bercovich , Edgar Altszyler , Juan Manuel Pérez

We introduce the Self-Annotated Reddit Corpus (SARC), a large corpus for sarcasm research and for training and evaluating systems for sarcasm detection. The corpus has 1.3 million sarcastic statements -- 10 times more than any previous…

计算与语言 · 计算机科学 2018-03-26 Mikhail Khodak , Nikunj Saunshi , Kiran Vodrahalli

Neural text generation (data- or text-to-text) demonstrates remarkable performance when training data is abundant which for many applications is not the case. To collect a large corpus of parallel data, heuristic rules are often used but…

计算与语言 · 计算机科学 2020-10-13 Katja Filippova

Clinical text provides essential information to estimate the acuity of a patient during hospital stays in addition to structured clinical data. In this study, we explore how clinical text can complement a clinical predictive learning task.…

The utilization of clinical reports for various secondary purposes, including health research and treatment monitoring, is crucial for enhancing patient care. Natural Language Processing (NLP) tools have emerged as valuable assets for…

计算与语言 · 计算机科学 2023-06-14 Iker de la Iglesia , Aitziber Atutxa , Koldo Gojenola , Ander Barrena

This paper introduces the Ubuntu Dialogue Corpus, a dataset containing almost 1 million multi-turn dialogues, with a total of over 7 million utterances and 100 million words. This provides a unique resource for research into building…

计算与语言 · 计算机科学 2016-07-26 Ryan Lowe , Nissan Pow , Iulian Serban , Joelle Pineau

In the field of image classification, existing methods often struggle with biased or ambiguous data, a prevalent issue in real-world scenarios. Current strategies, including semi-supervised learning and class blending, offer partial…

计算机视觉与模式识别 · 计算机科学 2024-04-30 Lars Schmarje , Vasco Grossmann , Claudius Zelenka , Johannes Brünger , Reinhard Koch

Understanding entailment and contradiction is fundamental to understanding natural language, and inference about entailment and contradiction is a valuable testing ground for the development of semantic representations. However, machine…

计算与语言 · 计算机科学 2015-08-24 Samuel R. Bowman , Gabor Angeli , Christopher Potts , Christopher D. Manning

Some of the major limitations identified in the areas of argument mining, argument generation, and natural language argument analysis are related to the complexity of annotating argumentatively rich data, the limited size of these corpora,…

计算与语言 · 计算机科学 2024-02-23 Ramon Ruiz-Dolz , Joaquin Taverner , John Lawrence , Chris Reed

Deployed language models must decide not only what to answer but also when not to answer. We present UniCR, a unified framework that turns heterogeneous uncertainty evidence including sequence likelihoods, self-consistency dispersion,…

A reasonable approach for fact checking a claim involves retrieving potentially relevant documents from different sources (e.g., news websites, social media, etc.), determining the stance of each document with respect to the claim, and…

计算与语言 · 计算机科学 2018-04-24 Ramy Baly , Mitra Mohtarami , James Glass , Lluis Marquez , Alessandro Moschitti , Preslav Nakov

Healthcare foundation models have largely followed paradigms from natural language processing and computer vision, emphasizing large scale pretraining and deterministic representations over heterogeneous clinical data. However, clinical…

机器学习 · 计算机科学 2026-04-07 Qian Zhou , Yuanyun Zhang , Shi Li

Negation is a fundamental linguistic concept used by humans to convey information that they do not desire. Despite this, minimal research has focused on negation within text-guided image editing. This lack of research means that…

计算机视觉与模式识别 · 计算机科学 2025-04-08 Nhat-Tan Bui , Dinh-Hieu Hoang , Quoc-Huy Trinh , Minh-Triet Tran , Truong Nguyen , Susan Gauch

The rapidly growing volume of scientific publications offers an interesting challenge for research on methods for analyzing the authorship of documents with one or more authors. However, most existing datasets lack scientific documents or…

计算与语言 · 计算机科学 2023-05-11 Janek Bevendorff , Philipp Sauer , Lukas Gienapp , Wolfgang Kircheis , Erik Körner , Benno Stein , Martin Potthast

We present a dataset of 833k paragraphs extracted from CC-BY licensed scientific publications, classified into four categories: acknowledgments, data mentions, software/code mentions, and clinical trial mentions. The paragraphs are…

计算与语言 · 计算机科学 2025-10-28 Eric Jeangirard

We describe a gold standard corpus of protest events that comprise of various local and international sources from various countries in English. The corpus contains document, sentence, and token level annotations. This corpus facilitates…

计算与语言 · 计算机科学 2020-08-04 Ali Hürriyetoğlu , Erdem Yörük , Deniz Yüret , Osman Mutlu , Çağrı Yoltar , Fırat Duruşan , Burak Gürel