English
Related papers

Related papers: MegaWika: Millions of reports and their sources ac…

200 papers

While Wikipedia exists in 287 languages, its content is unevenly distributed among them. In this work, we investigate the generation of open domain Wikipedia summaries in underserved languages using structured data from Wikidata. To this…

Computation and Language · Computer Science 2018-05-01 Lucie-Aimée Kaffee , Hady Elsahar , Pavlos Vougiouklis , Christophe Gravier , Frédérique Laforest , Jonathon Hare , Elena Simperl

To cope with the large number of publications, more and more researchers are automatically extracting data of interest using natural language processing methods based on supervised learning. Much data, especially in the natural and…

Computation and Language · Computer Science 2025-03-19 Jan Göpfert , Patrick Kuckertz , Jann M. Weinand , Detlef Stolten

Wikipedia is the world's largest online encyclopedia, but maintaining article quality through collaboration is challenging. Wikipedia designed a quality scale, but with such a manual assessment process, many articles remain unassessed. We…

Computation and Language · Computer Science 2023-10-04 Pedro Miguel Moás , Carla Teixeira Lopes

As free online encyclopedias with massive volumes of content, Wikipedia and Wikidata are key to many Natural Language Processing (NLP) tasks, such as information retrieval, knowledge base building, machine translation, text classification,…

We release a corpus of 43 million atomic edits across 8 languages. These edits are mined from Wikipedia edit history and consist of instances in which a human editor has inserted a single contiguous phrase into, or deleted a single…

Computation and Language · Computer Science 2018-08-29 Manaal Faruqui , Ellie Pavlick , Ian Tenney , Dipanjan Das

In this paper we present the Wikipedia Cultural Diversity dataset. For each existing Wikipedia language edition, the dataset contains a classification of the articles that represent its associated cultural context, i.e. all concepts and…

Computers and Society · Computer Science 2019-06-11 Marc Miquel-Ribé , David Laniado

This paper presents a new way to increase interconnectivity in small Wikipedias (fewer than a 100,000 articles), by automatically linking articles based on interlanguage links. Many small Wikipedias have many articles with very few links,…

Social and Information Networks · Computer Science 2017-01-10 Michael Lotkowski

Nowadays, information describing navigation behaviour of internet users are used in several fields, e-commerce, economy, sociology and data science. Such information can be extracted from different knowledge bases, including…

Social and Information Networks · Computer Science 2020-08-18 Célestin Coquidé , Włodzimierz Lewoniewski

News article revision histories have the potential to give us novel insights across varied fields of linguistics and social sciences. In this work, we present, to our knowledge, the first publicly available dataset of news article revision…

Computation and Language · Computer Science 2022-07-01 Alexander Spangher , Jonathan May

We present an approach based on multilingual sentence embeddings to automatically extract parallel sentences from the content of Wikipedia articles in 85 languages, including several dialects or low-resource languages. We do not limit the…

Computation and Language · Computer Science 2019-07-17 Holger Schwenk , Vishrav Chaudhary , Shuo Sun , Hongyu Gong , Francisco Guzmán

Large language models hallucinate factual claims and struggle to ground their outputs in retrievable evidence, particularly in non-English languages. Existing resources impose a trade-off: structured knowledge bases lack textual grounding,…

Computation and Language · Computer Science 2026-05-15 Yingli Shen , Wen Lai , Jie Zhou , Xueren Zhang , Yudong Wang , Kangyang Luo , Shuo Wang , Ge Gao , Alexander Fraser , Maosong Sun

English Wikipedia has long been an important data source for much research and natural language machine learning modeling. The growth of non-English language editions of Wikipedia, greater computational resources, and calls for equity in…

Computers and Society · Computer Science 2022-04-07 Isaac Johnson , Emily Lescak

Acknowledged as one of the most successful online cooperative projects in human society, Wikipedia has obtained rapid growth in recent years and desires continuously to expand content and disseminate knowledge values for everyone globally.…

Computation and Language · Computer Science 2022-10-25 Hoang Thang Ta , Alexander Gelbukha , Grigori Sidorov

The increasing diversity of languages used on the web introduces a new level of complexity to Information Retrieval (IR) systems. We can no longer assume that textual content is written in one language or even the same language family. In…

Computation and Language · Computer Science 2014-10-15 Rami Al-Rfou , Vivek Kulkarni , Bryan Perozzi , Steven Skiena

The Speech Wikimedia Dataset is a publicly available compilation of audio with transcriptions extracted from Wikimedia Commons. It includes 1780 hours (195 GB) of CC-BY-SA licensed transcribed speech from a diverse set of scenarios and…

Artificial Intelligence · Computer Science 2023-08-31 Rafael Mosquera Gómez , Julián Eusse , Juan Ciro , Daniel Galvez , Ryan Hileman , Kurt Bollacker , David Kanter

Timeline generation is of great significance for a comprehensive understanding of the development of events over time. Its goal is to organize news chronologically, which helps to identify patterns and trends that may be obscured when…

Information Retrieval · Computer Science 2025-02-12 Xiaochen Liu , Yanan Zhang

We present V\=arta, a large-scale multilingual dataset for headline generation in Indic languages. This dataset includes 41.8 million news articles in 14 different Indic languages (and English), which come from a variety of high-quality…

Computation and Language · Computer Science 2023-05-11 Rahul Aralikatte , Ziling Cheng , Sumanth Doddapaneni , Jackie Chi Kit Cheung

In this article we address the problem of text passage alignment across interlingual article pairs in Wikipedia. We develop methods that enable the identification and interlinking of text passages written in different languages and…

Computation and Language · Computer Science 2019-05-22 Simon Gottschalk , Elena Demidova

In the past decade, the DBpedia community has put significant amount of effort on developing technical infrastructure and methods for efficient extraction of structured information from Wikipedia. These efforts have been primarily focused…

Computation and Language · Computer Science 2018-12-27 Milan Dojchinovski , Julio Hernandez , Markus Ackermann , Amit Kirschenbaum , Sebastian Hellmann

Knowledge discovery and collection are intelligence-intensive tasks that traditionally require significant human effort to ensure high-quality outputs. Recent research has explored multi-agent frameworks for automating Wikipedia-style…

Computer Vision and Pattern Recognition · Computer Science 2025-09-08 Zhongyu Yang , Jun Chen , Dannong Xu , Junjie Fei , Xiaoqian Shen , Liangbing Zhao , Chun-Mei Feng , Mohamed Elhoseiny