English
Related papers

Related papers: How Grounded is Wikipedia? A Study on Structured E…

200 papers

Wikipedia is the world's largest online encyclopedia, but maintaining article quality through collaboration is challenging. Wikipedia designed a quality scale, but with such a manual assessment process, many articles remain unassessed. We…

Computation and Language · Computer Science 2023-10-04 Pedro Miguel Moás , Carla Teixeira Lopes

Wikipedia, as a social phenomenon of collaborative knowledge creating, has been studied extensively from various points of views. The category system of Wikipedia, introduced in 2004, has attracted relatively little attention. In this…

Physics and Society · Physics 2012-03-06 Krzysztof Suchecki , Alkim Almila Akdag Salah , Cheng Gao , Andrea Scharnhorst

We present a new dataset of Wikipedia articles each paired with a knowledge graph, to facilitate the research in conditional text generation, graph generation and graph representation learning. Existing graph-text paired datasets typically…

Computation and Language · Computer Science 2021-07-21 Luyu Wang , Yujia Li , Ozlem Aslan , Oriol Vinyals

In this paper we address the challenge of assessing the quality of Wikipedia pages using scores derived from edit contribution and contributor authoritativeness measures. The hypothesis is that pages with significant contributions from…

Social and Information Networks · Computer Science 2013-10-25 Xiangju Qin , Pádraig Cunningham

We present a study on predicting the factuality of reporting and bias of news media. While previous work has focused on studying the veracity of claims or documents, here we are interested in characterizing entire news media. These are…

Information Retrieval · Computer Science 2018-10-04 Ramy Baly , Georgi Karadzhov , Dimitar Alexandrov , James Glass , Preslav Nakov

Social media platforms, increasingly used as news sources for varied data analytics, have transformed how information is generated and disseminated. However, the unverified nature of this content raises concerns about trustworthiness and…

Information Retrieval · Computer Science 2025-03-10 Francisco de Arriba-Pérez , Silvia García-Méndez , Fátima Leal , Benedita Malheiro , Juan C Burguillo

Fast-developing fields such as Artificial Intelligence (AI) often outpace the efforts of encyclopedic sources such as Wikipedia, which either do not completely cover recently-introduced topics or lack such content entirely. As a result,…

Computation and Language · Computer Science 2022-06-23 Irene Li , Alexander Fabbri , Rina Kawamura , Yixin Liu , Xiangru Tang , Jaesung Tae , Chang Shen , Sally Ma , Tomoe Mizutani , Dragomir Radev

Peer production platforms like Wikipedia commonly suffer from content gaps. Prior research suggests recommender systems can help solve this problem, by guiding editors towards underrepresented topics. However, it remains unclear whether…

Computers and Society · Computer Science 2024-04-11 Mo Houtti , Isaac Johnson , Morten Warncke-Wang , Loren Terveen

With 60M articles in more than 300 language versions, Wikipedia is the largest platform for open and freely accessible knowledge. While the available content has been growing continuously at a rate of around 200K new articles each month,…

Social and Information Networks · Computer Science 2024-10-08 Akhil Arora , Robert West , Martin Gerlach

The state-of-the-art named entity recognition (NER) systems are statistical machine learning models that have strong generalization capability (i.e., can recognize unseen entities that do not appear in training data) based on lexical and…

Computation and Language · Computer Science 2019-11-04 Jian Ni , Radu Florian

Wikipedia has been turned into an immensely popular crowd-sourced encyclopedia for information dissemination on numerous versatile topics in the form of subscription free content. It allows anyone to contribute so that the articles remain…

Social and Information Networks · Computer Science 2021-11-03 Paramita Das , Bhanu Prakash Reddy Guda , Sasi Bhusan Seelaboyina , Soumya Sarkar , Animesh Mukherjee

Knowledge bases are collections of domain-specific and commonsense facts. Recently, the sizes of KBs are rocketing due to automatic extraction for knowledge and facts. For example, the number of facts in WikiData is up to 974 million!…

Databases · Computer Science 2021-10-22 Ruoyu Wang , Daniel Sun , Guoqiang Li , Raymond Wong , Shiping Chen

We study how to apply large language models to write grounded and organized long-form articles from scratch, with comparable breadth and depth to Wikipedia pages. This underexplored problem poses new challenges at the pre-writing stage,…

Computation and Language · Computer Science 2024-04-09 Yijia Shao , Yucheng Jiang , Theodore A. Kanell , Peter Xu , Omar Khattab , Monica S. Lam

Clarifying the research framing of NLP artefacts (e.g., models, datasets, etc.) is crucial to aligning research with practical applications. Recent studies manually analyzed NLP research across domains, showing that few papers explicitly…

Computation and Language · Computer Science 2025-10-07 Eric Chamoun , Nedjma Ousidhoum , Michael Schlichtkrull , Andreas Vlachos

The dataset focuses on Wikipedia users and contains information about demographic and socioeconomic characteristics of the respondents and their activity on Wikipedia. The data was collected using a questionnaire available online between…

Computers and Society · Computer Science 2023-12-06 Caterina Cruciani , Léo Joubert , Nicolas Jullien , Laurent Mell , Sasha Piccione , Jeanne Vermeirsche

In this paper, we describe an embedding-based entity recommendation framework for Wikipedia that organizes Wikipedia into a collection of graphs layered on top of each other, learns complementary entity representations from their topology…

Information Retrieval · Computer Science 2020-04-16 Chien-Chun Ni , Kin Sum Liu , Nicolas Torzec

This paper was originally designed as a literature review for a doctoral dissertation focusing on Wikipedia. This exposition gives the structure of Wikipedia and the latest trends in Wikipedia research.

Digital Libraries · Computer Science 2015-03-19 Owen S. Martin

Social norms have traditionally been difficult to quantify. In any particular society, their sheer number and complex interdependencies often limit a system-level analysis. One exception is that of the network of norms that sustain the…

Social and Information Networks · Computer Science 2016-05-19 Bradi Heaberlin , Simon DeDeo

Wikipedia is the largest source of free encyclopedic knowledge and one of the most visited sites on the Web. To increase reader understanding of the article, Wikipedia editors add images within the text of the article's body. However,…

Computers and Society · Computer Science 2021-12-06 Daniele Rama , Tiziano Piccardi , Miriam Redi , Rossano Schifanella

In this paper, we present a dataset of inter-language knowledge propagation in Wikipedia. Covering the entire 309 language editions and 33M articles, the dataset aims to track the full propagation history of Wikipedia concepts, and allow…

Computers and Society · Computer Science 2021-04-01 Roldolfo Valentim , Giovanni Comarela , Souneil Park , Diego Saez-Trumper