中文
相关论文

相关论文: Wiki Dumps to Training Corpora: South Slavic Case

200 篇论文

Split and rephrase is the task of breaking down a sentence into shorter ones that together convey the same meaning. We extract a rich new dataset for this task by mining Wikipedia's edit history: WikiSplit contains one million naturally…

计算与语言 · 计算机科学 2018-08-30 Jan A. Botha , Manaal Faruqui , John Alex , Jason Baldridge , Dipanjan Das

Grammatical Error Correction (GEC) has been recently modeled using the sequence-to-sequence framework. However, unlike sequence transduction problems such as machine translation, GEC suffers from the lack of plentiful parallel data. We…

计算与语言 · 计算机科学 2019-04-12 Jared Lichtarge , Chris Alberti , Shankar Kumar , Noam Shazeer , Niki Parmar , Simon Tong

We present an approach based on multilingual sentence embeddings to automatically extract parallel sentences from the content of Wikipedia articles in 85 languages, including several dialects or low-resource languages. We do not limit the…

计算与语言 · 计算机科学 2019-07-17 Holger Schwenk , Vishrav Chaudhary , Shuo Sun , Hongyu Gong , Francisco Guzmán

There are large amounts of insight and social discovery potential in mining crowd-sourced comments left on popular news forums like Reddit.com, Tumblr.com, Facebook.com and Hacker News. Unfortunately, due the overwhelming amount of…

计算与语言 · 计算机科学 2017-01-13 Manuel Amunategui

This paper gives comprehensive analyses of corpora based on Wikipedia for several tasks in question answering. Four recent corpora are collected,WikiQA, SelQA, SQuAD, and InfoQA, and first analyzed intrinsically by contextual similarities,…

计算与语言 · 计算机科学 2018-02-06 Tomasz Jurczyk , Amit Deshmane , Jinho D. Choi

Pre-training text representations have led to significant improvements in many areas of natural language processing. The quality of these models benefits greatly from the size of the pretraining corpora as long as its quality is preserved.…

We present a simple but effective approach for leveraging Wikipedia for neural machine translation as well as cross-lingual tasks of image captioning and dependency parsing without using any direct supervision from external parallel data or…

计算与语言 · 计算机科学 2021-09-13 Mohammad Sadegh Rasooli , Chris Callison-Burch , Derry Tanti Wijaya

The use of domain knowledge is generally found to improve query efficiency in content filtering applications. In particular, tangible benefits have been achieved when using knowledge-based approaches within more specialized fields, such as…

信息检索 · 计算机科学 2015-03-17 Pekka Malo , Pyry Siitari , Oskar Ahlgren , Jyrki Wallenius , Pekka Korhonen

Text simplification research has mostly focused on sentence-level simplification, even though many desirable edits - such as adding relevant background information or reordering content - may require document-level context. Prior work has…

计算与语言 · 计算机科学 2023-05-31 Philippe Laban , Jesse Vig , Wojciech Kryscinski , Shafiq Joty , Caiming Xiong , Chien-Sheng Wu

We study the problem of entity salience by proposing the design and implementation of SWAT, a system that identifies the salient Wikipedia entities occurring in an input document. SWAT consists of several modules that are able to detect and…

信息检索 · 计算机科学 2019-05-17 Marco Ponza , Paolo Ferragina , Francesco Piccinno

Understanding procedural natural language (e.g., step-by-step instructions) is a crucial step to execution and planning. However, while there are ample corpora and downstream tasks available in English, the field lacks such resources for…

计算与语言 · 计算机科学 2024-03-08 Arda Uzunoglu , Gözde Gül Şahin

Wikidata is one of the most important sources of structured data on the web, built by a worldwide community of volunteers. As a secondary source, its contents must be backed by credible references; this is particularly important as Wikidata…

人工智能 · 计算机科学 2021-09-21 Gabriel Amaral , Alessandro Piscopo , Lucie-Aimée Kaffee , Odinaldo Rodrigues , Elena Simperl

We introduce a next-generation vandalism detection system for Wikidata, one of the largest open-source structured knowledge bases on the Web. Wikidata is highly complex: its items incorporate an ever-expanding universe of factual triples…

计算与语言 · 计算机科学 2025-05-26 Mykola Trokhymovych , Lydia Pintscher , Ricardo Baeza-Yates , Diego Saez-Trumper

We present a cross-lingual summarisation corpus with long documents in a source language associated with multi-sentence summaries in a target language. The corpus covers twelve language pairs and directions for four European languages,…

计算与语言 · 计算机科学 2022-02-22 Laura Perez-Beltrachini , Mirella Lapata

Prior work on Data-To-Text Generation, the task of converting knowledge graph (KG) triples into natural text, focused on domain-specific benchmark datasets. In this paper, however, we verbalize the entire English Wikidata KG, and discuss…

计算与语言 · 计算机科学 2021-03-16 Oshin Agarwal , Heming Ge , Siamak Shakeri , Rami Al-Rfou

In this work, we propose an automatic evaluation and comparison of the browsing behavior of Wikipedia readers that can be applied to any language editions of Wikipedia. As an example, we focus on English, French, and Russian languages…

社会与信息网络 · 计算机科学 2020-02-18 Volodymyr Miz , Joëlle Hanna , Nicolas Aspert , Benjamin Ricaud , Pierre Vandergheynst

In the recent years, transformer-based models have lead to significant advances in language modelling for natural language processing. However, they require a vast amount of data to be (pre-)trained and there is a lack of corpora in…

Datasets for data-to-text generation typically focus either on multi-domain, single-sentence generation or on single-domain, long-form generation. In this work, we cast generating Wikipedia sections as a data-to-text generation task and…

计算与语言 · 计算机科学 2021-06-03 Mingda Chen , Sam Wiseman , Kevin Gimpel

Wiki articles are created and maintained by a crowd of editors, producing a continuous stream of reviews. Reviews can take the form of additions, reverts, or both. This crowdsourcing model is exposed to manipulation since neither reviews…

计算与语言 · 计算机科学 2024-05-29 Silvia García Méndez , Fátima Leal , Benedita Malheiro , Juan Carlos Burguillo Rial

We study how to apply large language models to write grounded and organized long-form articles from scratch, with comparable breadth and depth to Wikipedia pages. This underexplored problem poses new challenges at the pre-writing stage,…

计算与语言 · 计算机科学 2024-04-09 Yijia Shao , Yucheng Jiang , Theodore A. Kanell , Peter Xu , Omar Khattab , Monica S. Lam