English
Related papers

Related papers: Investigating Deletion in Wikipedia

200 papers

Online IR tools have to take into account new phenomena linked to the appearance of blogs, wiki and other collaborative publications. Among these collaborative sites, Wikipedia represents a crucial source of information. However, the…

Information Retrieval · Computer Science 2008-12-18 Bernard Jacquemin , Aurélien Lauf , Céline Poudat , Martine Hurault-Plantet , Nicolas Auray

Among the manifold takes on world literature, it is our goal to contribute to the discussion from a digital point of view by analyzing the representation of world literature in Wikipedia with its millions of articles in hundreds of…

Information Retrieval · Computer Science 2017-01-05 Christoph Hube , Frank Fischer , Robert Jäschke , Gerhard Lauer , Mads Rosendahl Thomsen

"Keyword Extraction" refers to the task of automatically identifying the most relevant and informative phrases in natural language text. As we are deluged with large amounts of text data in many different forms and content - emails, blogs,…

Computation and Language · Computer Science 2019-08-22 Shibamouli Lahiri

Over 500 million tweets are posted in Twitter each day, out of which about 11% tweets are deleted by the users posting them. This phenomenon of widespread deletion of tweets leads to a number of questions: what kind of content posted by…

Social and Information Networks · Computer Science 2022-12-27 Parantapa Bhattacharya , Saptarshi Ghosh , Niloy Ganguly

In recent years, there has been a huge increase in the number of bots online, varying from Web crawlers for search engines, to chatbots for online customer service, spambots on social media, and content-editing bots in online collaboration…

Social and Information Networks · Computer Science 2017-02-28 Milena Tsvetkova , Ruth García-Gavilanes , Luciano Floridi , Taha Yasseri

Wikipedia is the biggest encyclopedia ever created and the fifth most visited website in the world. Tens of millions of people surf it every day, seeking answers to various questions. Collective user activity on its pages leaves publicly…

Information Retrieval · Computer Science 2018-02-15 Volodymyr Miz , Kirell Benzi , Benjamin Ricaud , Pierre Vandergheynst

Looking into the growth of information in the web it is a very tedious process of getting the exact information the user is looking for. Many search engines generate user profile related data listing. This paper involves one such process…

Information Retrieval · Computer Science 2011-09-12 L. K. Joshila Grace , V. Maheswari , Dhinaharan Nagamalai

With this work, we present a publicly available dataset of the history of all the references (more than 55 million) ever used in the English Wikipedia until June 2019. We have applied a new method for identifying and monitoring references…

Computers and Society · Computer Science 2020-10-08 Olga Zagovora , Roberto Ulloa , Katrin Weller , Fabian Flöck

News article revision histories provide clues to narrative and factual evolution in news articles. To facilitate analysis of this evolution, we present the first publicly available dataset of news revision histories, NewsEdits. Our dataset…

Computation and Language · Computer Science 2022-06-16 Alexander Spangher , Xiang Ren , Jonathan May , Nanyun Peng

The Library of Babel, described by Jorge Luis Borges, stores an enormous amount of information. The Library exists {\it ab aeterno}. Wikipedia, a free online encyclopaedia, becomes a modern analogue of such a Library. Information retrieval…

Information Retrieval · Computer Science 2010-11-15 A. O. Zhirov , O. V. Zhirov , D. L. Shepelyansky

Hyperlinks and other relations in Wikipedia are a extraordinary resource which is still not fully understood. In this paper we study the different types of links in Wikipedia, and contrast the use of the full graph with respect to just…

Computation and Language · Computer Science 2015-03-16 Eneko Agirre , Ander Barrena , Aitor Soroa

Wikipedia is one of the richest knowledge sources on the Web today. In order to facilitate navigating, searching, and maintaining its content, Wikipedia's guidelines state that all articles should be annotated with a so-called short…

Computation and Language · Computer Science 2023-02-20 Marija Sakota , Maxime Peyrard , Robert West

We make decisions by reacting to changes in the real world, in particular, the emergence and disappearance of impermanent entities such as events, restaurants, and services. Because we want to avoid missing out on opportunities or making…

Computation and Language · Computer Science 2022-10-17 Satoshi Akasaki , Naoki Yoshinaga , Masashi Toyoda

Wikipedia is a huge opportunity for machine learning, being the largest semi-structured base of knowledge available. Because of this, many works examine its contents, and focus on structuring it in order to make it usable in learning tasks,…

Machine Learning · Computer Science 2020-01-23 Tiphaine Viard , Thomas McLachlan , Hamidreza Ghader , Satoshi Sekine

Statistics of article page views is useful for measuring the impact of individual articles. Analyzing the temporal evolution of article page views, we find that article page views usually decay over time after reaching a peak, especially…

Physics and Society · Physics 2016-06-17 Yeseul Kim , Kun Cho , Byung Mook Weon

We introduce MegaWika 2, a large, multilingual dataset of Wikipedia articles with their citations and scraped web sources; articles are represented in a rich data structure, and scraped source texts are stored inline with precise character…

Digital Libraries · Computer Science 2025-08-07 Samuel Barham , Chandler May , Benjamin Van Durme

English Wikipedia has long been an important data source for much research and natural language machine learning modeling. The growth of non-English language editions of Wikipedia, greater computational resources, and calls for equity in…

Computers and Society · Computer Science 2022-04-07 Isaac Johnson , Emily Lescak

The present study aims to establish a valid method by which to apply the theory of co-citations to Wikipedia article references and, subsequently, to map these relationships between scientific papers. This theory, originally applied to…

Digital Libraries · Computer Science 2019-07-31 Daniel Torres-Salinas , Esteban Romero-Frías , Wenceslao Arroyo-Machado

The cumulative effect of collective online participation has an important and adverse impact on individual privacy. As an online system evolves over time, new digital traces of individual behavior may uncover previously hidden statistical…

Social and Information Networks · Computer Science 2015-12-18 Marian-Andrei Rizoiu , Lexing Xie , Tiberio Caetano , Manuel Cebrian

The Web has drastically simplified our access to knowledge and learning, and fact-checking online resources has become a part of our daily routine. Studying online knowledge consumption is thus critical for understanding human behavior and…

Computers and Society · Computer Science 2025-01-03 Tiziano Piccardi , Robert West