English
Related papers

Related papers: Building automated vandalism detection tools for W…

200 papers

Wikipedia is nowadays a widely used encyclopedia, and one of the most visible sites on the Internet. Its strong principle of collaborative work and free editing sometimes generates disputes due to disagreements between users. In this…

Information Retrieval · Computer Science 2008-12-18 Bernard Jacquemin , Aurélien Lauf , Céline Poudat , Martine Hurault-Plantet , Nicolas Auray

Wikipedia categories, a classification scheme built for organizing and describing Wikpedia articles, are being applied in computer science research. This paper adopts a systematic literature review approach, in order to identify different…

Digital Libraries · Computer Science 2020-04-22 Jesús Tramullas , Piedad Garrido-Picazo , Ana I. Sánchez-Casabón

This paper proposes to tackle open- domain question answering using Wikipedia as the unique knowledge source: the answer to any factoid question is a text span in a Wikipedia article. This task of machine reading at scale combines the…

Computation and Language · Computer Science 2017-05-01 Danqi Chen , Adam Fisch , Jason Weston , Antoine Bordes

We present a new concept - Wikiometrics - the derivation of metrics and indicators from Wikipedia. Wikipedia provides an accurate representation of the real world due to its size, structure, editing policy and popularity. We demonstrate an…

Digital Libraries · Computer Science 2016-01-11 Gilad Katz , Lior Rokach

We describe NatCat, a large-scale resource for text classification constructed from three data sources: Wikipedia, Stack Exchange, and Reddit. NatCat consists of document-category pairs derived from manual curation that occurs naturally…

Computation and Language · Computer Science 2021-09-21 Zewei Chu , Karl Stratos , Kevin Gimpel

This paper presents an in-depth analysis of Wikidata qualifiers, focusing on their semantics and actual usage, with the aim of developing a taxonomy that addresses the challenges of selecting appropriate qualifiers, querying the graph, and…

Artificial Intelligence · Computer Science 2026-03-13 Gilles Falquet , Sahar Aljalbout

Wikipedia is the biggest encyclopedia ever created and the fifth most visited website in the world. Tens of millions of people surf it every day, seeking answers to various questions. Collective user activity on its pages leaves publicly…

Information Retrieval · Computer Science 2018-02-15 Volodymyr Miz , Kirell Benzi , Benjamin Ricaud , Pierre Vandergheynst

Plagiarism is an act of using someone else's work without proper acknowledgment, and this sin is seen to cut across various arenas including the academy, publishing, and other similar arenas. The traditional methods of plagiarism detection…

Emerging Technologies · Computer Science 2024-12-10 Omraj Kamat , Tridib Ghosh , Kalaivani J , Angayarkanni V , Rama P

The Internet-based encyclopaedia Wikipedia has grown to become one of the most visited web-sites on the Internet. However, critics have questioned the quality of entries, and an empirical study has shown Wikipedia to contain errors in a…

Digital Libraries · Computer Science 2011-01-04 Finn Aarup Nielsen

Social media plays a significant role in sharing essential information, which helps humanitarian organizations in rescue operations during and after disaster incidents. However, developing an efficient method that can provide rapid analysis…

Computer Vision and Pattern Recognition · Computer Science 2022-02-02 Soroor Shekarizadeh , Razieh Rastgoo , Saif Al-Kuwari , Mohammad Sabokrou

Socialization in online communities allows existing members to welcome and recruit newcomers, introduce them to community norms and practices, and sustain their early participation. However, socializing newcomers does not come for free: in…

Social and Information Networks · Computer Science 2014-10-14 Giovanni Luca Ciampaglia , Dario Taraborelli

The use of domain knowledge is generally found to improve query efficiency in content filtering applications. In particular, tangible benefits have been achieved when using knowledge-based approaches within more specialized fields, such as…

Information Retrieval · Computer Science 2015-03-17 Pekka Malo , Pyry Siitari , Oskar Ahlgren , Jyrki Wallenius , Pekka Korhonen

Algorithmic systems---from rule-based bots to machine learning classifiers---have a long history of supporting the essential work of content moderation and other curation work in peer production projects. From counter-vandalism to task…

Human-Computer Interaction · Computer Science 2020-08-21 Aaron Halfaker , R. Stuart Geiger

Knowledge bases such as Wikidata, DBpedia, or YAGO contain millions of entities and facts. In some knowledge bases, the correctness of these facts has been evaluated. However, much less is known about their completeness, i.e., the…

Databases · Computer Science 2016-12-20 Luis Galárraga , Simon Razniewski , Antoine Amarilli , Fabian M. Suchanek

Several hundred Wikipedia articles are deleted every day because they lack sufficient significance to be included in the encyclopedia. We collect a dataset of deleted articles and analyze them to determine whether or not the deletions were…

Computers and Society · Computer Science 2013-05-24 Bluma S. Gelley

This paper is an introductory discussion on the cause of open source software vulnerabilities, their importance in the cybersecurity ecosystem, and a selection of detection methods. A recent application security report showed 44% of…

Cryptography and Security · Computer Science 2022-03-31 Stuart Millar

This paper presents the first attempt, up to our knowledge, to classify English writing styles on this scale with the challenge of classifying day to day language written by writers with different backgrounds covering various areas of…

Computation and Language · Computer Science 2017-04-26 Yanging Chen , Rami Al-Rfou' , Yejin Choi

Wikipedia's perceived high quality and broad language coverage have established it as a fundamental resource in NLP. However, in recent years, such assumptions of high quality have become the subject of scrutiny in low-resource and…

In this paper we present the Wikipedia Cultural Diversity dataset. For each existing Wikipedia language edition, the dataset contains a classification of the articles that represent its associated cultural context, i.e. all concepts and…

Computers and Society · Computer Science 2019-06-11 Marc Miquel-Ribé , David Laniado

For more than a decade now, academicians and online platform administrators have been studying solutions to the problem of bot detection. Bots are computer algorithms whose use is far from being benign: malicious bots are purposely created…

Cryptography and Security · Computer Science 2025-06-25 Rocco De Nicola , Marinella Petrocchi , Manuel Pratelli