中文
相关论文

相关论文: A Corpus-Based Investigation of Definite Descripti…

200 篇论文

Large sense-annotated datasets are increasingly necessary for training deep supervised systems in Word Sense Disambiguation. However, gathering high-quality sense-annotated data for as many instances as possible is a laborious and expensive…

计算与语言 · 计算机科学 2020-03-16 Tommaso Pasini , Jose Camacho-Collados

Quotation extraction is a widely useful task both from a sociological and from a Natural Language Processing perspective. However, very little data is available to study this task in languages other than English. In this paper, we present a…

计算与语言 · 计算机科学 2023-09-20 Ange Richard , Laura Alonzo-Canul , François Portet

Modeling thematic fit (a verb--argument compositional semantics task) currently requires a very large burden of labeled data. We take a linguistically machine-annotated large corpus and replace corpus layers with output from higher-quality,…

计算与语言 · 计算机科学 2022-05-05 Yuval Marton , Asad Sayeed

Temporal relation extraction models have thus far been hindered by a number of issues in existing temporal relation-annotated news datasets, including: (1) low inter-annotator agreement due to the lack of specificity of their annotation…

计算与语言 · 计算机科学 2023-10-30 Sarah Alsayyahi , Riza Batista-Navarro

Paraphrase detection is important for a number of applications, including plagiarism detection, authorship attribution, question answering, text summarization, text mining in general, etc. In this paper, we give a performance overview of…

计算与语言 · 计算机科学 2021-06-02 Tedo Vrbanec , Ana Mestrovic

Identifying near duplicates within large, noisy text corpora has a myriad of applications that range from de-duplicating training datasets, reducing privacy risk, and evaluating test set leakage, to identifying reproduced news articles and…

计算与语言 · 计算机科学 2024-04-25 Emily Silcock , Luca D'Amico-Wong , Jinglin Yang , Melissa Dell

As natural language corpora expand at an unprecedented rate, manual annotation remains a significant methodological bottleneck in corpus linguistic work. We address this challenge by presenting a scalable pipeline for automating grammatical…

计算与语言 · 计算机科学 2026-02-11 Cameron Morin , Matti Marttinen Larsson

Rapid and efficient assessment of the future impact of research articles is a significant concern for both authors and reviewers. The most common standard for measuring the impact of academic papers is the number of citations. In recent…

计算与语言 · 计算机科学 2025-03-28 Qichen Sun , Yuxing Lu , Kun Xia , Li Chen , He Sun , Jinzhuo Wang

Attribute-based recognition models, due to their impressive performance and their ability to generalize well on novel categories, have been widely adopted for many computer vision applications. However, usually both the attribute vocabulary…

计算机视觉与模式识别 · 计算机科学 2017-04-13 Ziad Al-Halah , Rainer Stiefelhagen

The past few years have witnessed renewed interest in NLP tasks at the interface between vision and language. One intensively-studied problem is that of automatically generating text from images. In this paper, we extend this problem to the…

This paper presents the results of a study on the semantic constraints imposed on lexical choice by certain contextual indicators. We show how such indicators are computed and how correlations between them and the choice of a noun phrase…

cmp-lg · 计算机科学 2007-05-23 Dragomir R. Radev

We present a corpus of 5,000 richly annotated abstracts of medical articles describing clinical randomized controlled trials. Annotations include demarcations of text spans that describe the Patient population enrolled, the Interventions…

计算与语言 · 计算机科学 2018-06-13 Benjamin Nye , Junyi Jessy Li , Roma Patel , Yinfei Yang , Iain J. Marshall , Ani Nenkova , Byron C. Wallace

This paper presents a new challenging information extraction task in the domain of materials science. We develop an annotation scheme for marking information on experiments related to solid oxide fuel cells in scientific publications, such…

计算与语言 · 计算机科学 2020-06-05 Annemarie Friedrich , Heike Adel , Federico Tomazic , Johannes Hingerl , Renou Benteau , Anika Maruscyk , Lukas Lange

Coreference resolution is an important task for natural language understanding, and the resolution of ambiguous pronouns a longstanding challenge. Nonetheless, existing corpora do not capture ambiguous pronouns in sufficient volume or…

计算与语言 · 计算机科学 2018-10-15 Kellie Webster , Marta Recasens , Vera Axelrod , Jason Baldridge

This paper introduces the SAMSum Corpus, a new dataset with abstractive dialogue summaries. We investigate the challenges it poses for automated summarization by testing several models and comparing their results with those obtained on a…

计算与语言 · 计算机科学 2019-12-02 Bogdan Gliwa , Iwona Mochol , Maciej Biesek , Aleksander Wawer

Scholarly documents have a great degree of variation, both in terms of content (semantics) and structure (pragmatics). Prior work in scholarly document understanding emphasizes semantics through document summarization and corpus topic…

计算与语言 · 计算机科学 2023-10-03 Lee Kezar , Jay Pujara

In essence, embedding algorithms work by optimizing the distance between a word and its usual context in order to generate an embedding space that encodes the distributional representation of words. In addition to single words or word…

计算与语言 · 计算机科学 2021-04-14 Andres Garcia-Silva , Ronald Denaux , Jose Manuel Gomez-Perez

Text summarization models are approaching human levels of fidelity. Existing benchmarking corpora provide concordant pairs of full and abridged versions of Web, news or, professional content. To date, all summarization datasets operate…

计算与语言 · 计算机科学 2022-06-01 Seyed Ali Bahrainian , Sheridan Feucht , Carsten Eickhoff

Language models perform well on grammatical agreement, but it is unclear whether this reflects rule-based generalization or memorization. We study this question for German definite singular articles, whose forms depend on gender and case.…

计算与语言 · 计算机科学 2026-04-15 Jonathan Drechsel , Erisa Bytyqi , Steffen Herbold

In this paper, we compose a new task for deep argumentative structure analysis that goes beyond shallow discourse structure analysis. The idea is that argumentative relations can reasonably be represented with a small set of predefined…

计算与语言 · 计算机科学 2017-12-08 Paul Reisert , Naoya Inoue , Naoaki Okazaki , Kentaro Inui