中文
相关论文

相关论文: A Corpus-Based Investigation of Definite Descripti…

200 篇论文

Much linguistic research relies on annotated datasets of features extracted from text corpora, but the rapid quantitative growth of these corpora has created practical difficulties for linguists to manually annotate large data samples. In…

计算与语言 · 计算机科学 2025-04-11 Cameron Morin , Matti Marttinen Larsson

A variety of statistical methods for noun compound analysis are implemented and compared. The results support two main conclusions. First, the use of conceptual association not only enables a broad coverage, but also improves the accuracy.…

cmp-lg · 计算机科学 2008-02-03 Mark Lauer

This paper elaborates on the notion of uncertainty in the context of annotation in large text corpora, specifically focusing on (but not limited to) historical languages. Such uncertainty might be due to inherent properties of the language,…

计算与语言 · 计算机科学 2021-05-31 Marie-Luis Merten , Marcel Wever , Michaela Geierhos , Doris Tophinke , Eyke Hüllermeier

We introduce a new benchmark for coreference resolution and NLI, Knowref, that targets common-sense understanding and world knowledge. Previous coreference resolution tasks can largely be solved by exploiting the number and gender of the…

计算与语言 · 计算机科学 2019-06-17 Ali Emami , Paul Trichelair , Adam Trischler , Kaheer Suleman , Hannes Schulz , Jackie Chi Kit Cheung

We present a novel corpus of 445 human- and computer-generated documents, comprising about 27,000 clauses, annotated for semantic clause types and coherence relations that allow for nuanced comparison of artificial and natural discourse…

In this paper, we propose a new annotation scheme to classify different types of clauses in Terms-and-Conditions contracts with the ultimate goal of supporting legal experts to quickly identify and assess problematic issues in this type of…

计算与语言 · 计算机科学 2024-05-28 Pietro Giovanni Bizzaro , Elena Della Valentina , Maurizio Napolitano , Nadia Mana , Massimo Zancanaro

Discourse information is difficult to represent and annotate. Among the major frameworks for annotating discourse information, RST, PDTB and SDRT are widely discussed and used, each having its own theoretical foundation and focus. Corpora…

计算与语言 · 计算机科学 2022-04-19 Yingxue Fu

Analyzing how humans revise their writings is an interesting research question, not only from an educational perspective but also in terms of artificial intelligence. Better understanding of this process could facilitate many NLP…

计算与语言 · 计算机科学 2022-06-06 Omid Kashefi , Tazin Afrin , Meghan Dale , Christopher Olshefski , Amanda Godley , Diane Litman , Rebecca Hwa

Classic Topic Models are built under the Bag Of Words assumption, in which word position is ignored for simplicity. Besides, symmetric priors are typically used in most applications. In order to easily learn topics with different properties…

计算与语言 · 计算机科学 2018-06-27 Simón Roca-Sotelo , Jerónimo Arenas-García

In this paper, we present a new corpus of entailment problems. This corpus combines the following characteristics: 1. it is precise (does not leave out implicit hypotheses) 2. it is based on "real-world" texts (i.e. most of the premises…

计算与语言 · 计算机科学 2018-12-17 Jean-Philippe Bernardy , Stergios Chatzikyriakidis

Although pre-trained named entity recognition (NER) models are highly accurate on modern corpora, they underperform on historical texts due to differences in language OCR errors. In this work, we develop a new NER corpus of 3.6M sentences…

计算与语言 · 计算机科学 2023-06-08 Vít Novotný , Kristýna Luger , Michal Štefánik , Tereza Vrabcová , Aleš Horák

Summarizing texts is not a straightforward task. Before even considering text summarization, one should determine what kind of summary is expected. How much should the information be compressed? Is it relevant to reformulate or should the…

计算与语言 · 计算机科学 2020-07-16 Paul Tardy , David Janiszek , Yannick Estève , Vincent Nguyen

We describe an effort to annotate a corpus of natural language instructions consisting of 622 wet lab protocols to facilitate automatic or semi-automatic conversion of protocols into a machine-readable format and benefit biological…

计算与语言 · 计算机科学 2018-05-02 Chaitanya Kulkarni , Wei Xu , Alan Ritter , Raghu Machiraju

The acquisition of grammar has been a central question to adjudicate between theories of language acquisition. In order to conduct faster, more reproducible, and larger-scale corpus studies on grammaticality in child-caregiver…

计算与语言 · 计算机科学 2024-03-22 Mitja Nikolaus , Abhishek Agrawal , Petros Kaklamanis , Alex Warstadt , Abdellah Fourtassi

Scientific article summarization is challenging: large, annotated corpora are not available, and the summary should ideally include the article's impacts on research community. This paper provides novel solutions to these two challenges. We…

计算与语言 · 计算机科学 2019-09-17 Michihiro Yasunaga , Jungo Kasai , Rui Zhang , Alexander R. Fabbri , Irene Li , Dan Friedman , Dragomir R. Radev

Introduction: Clinical text classification using natural language processing (NLP) models requires adequate training data to achieve optimal performance. For that, 200-500 documents are typically annotated. The number is constrained by time…

This document, based on feedback from UMR TETIS members and the scientific literature, provides a generic methodology for creating annotation guidelines and annotated textual datasets (corpora). It covers methodological aspects, as well as…

信息检索 · 计算机科学 2026-01-21 Bahdja Boudoua , Nadia Guiffant , Mathieu Roche , Maguelonne Teisseire , Annelise Tran

Many communities, including the scientific community, develop implicit writing norms. Understanding them is crucial for effective communication with that community. Writers gradually develop an implicit understanding of norms by reading…

人机交互 · 计算机科学 2025-03-18 Hai Dang , Chelse Swoopes , Daniel Buschek , Elena L. Glassman

In terms of annotation structure, most learner corpora rely on holistic flat label inventories which, even when extensive, do not explicitly separate multiple linguistic dimensions. This makes linguistically deep annotation difficult and…

The corpus, from which a predictive language model is trained, can be considered the experience of a semantic system. We recorded everyday reading of two participants for two months on a tablet, generating individual corpus samples of…

计算与语言 · 计算机科学 2020-10-21 Markus J. Hofmann , Lara Müller , Andre Rölke , Ralph Radach , Chris Biemann