English
Related papers

Related papers: Large-Scale Vandalism Detection with Linear Classi…

200 papers

Wikidata is the new, large-scale knowledge base of the Wikimedia Foundation. As it can be edited by anyone, entries frequently get vandalized, leading to the possibility that it might spread of falsified information if such posts are not…

Information Retrieval · Computer Science 2017-12-20 Qi Zhu , Hongwei Ng , Liyuan Liu , Ziwei Ji , Bingjie Jiang , Jiaming Shen , Huan Gui

Wikidata is a free and open knowledge base from the Wikimedia Foundation, that not only acts as a central storage of structured data for other projects of the organization, but also for a growing array of information systems, including…

The WSDM Cup 2017 is a binary classification task for classifying Wikidata revisions into vandalism and non-vandalism. This paper describes our method using some machine learning techniques such as under-sampling, feature selection,…

Information Retrieval · Computer Science 2017-12-20 Tomoya Yamazaki , Mei Sasaki , Naoya Murakami , Takuya Makabe , Hiroki Iwasawa

We report on the Wikidata vandalism detection task at the WSDM Cup 2017. The task received five submissions for which this paper describes their evaluation and a comparison to state of the art baselines. Unlike previous work, we recast…

Information Retrieval · Computer Science 2017-12-19 Stefan Heindorf , Martin Potthast , Gregor Engels , Benno Stein

The WSDM Cup 2017 was a data mining challenge held in conjunction with the 10th International Conference on Web Search and Data Mining (WSDM). It addressed key challenges of knowledge bases today: quality assurance and entity search. For…

Information Retrieval · Computer Science 2017-12-29 Martin Potthast , Stefan Heindorf , Hannah Bast

Wikidata, like Wikipedia, is a knowledge base that anyone can edit. This open collaboration model is powerful in that it reduces barriers to participation and allows a large number of people to contribute. However, it exposes the knowledge…

Information Retrieval · Computer Science 2017-03-14 Amir Sarabadani , Aaron Halfaker , Dario Taraborelli

Wikipedia is an online encyclopedia that anyone can edit. In this open model, some people edits with the intent of harming the integrity of Wikipedia. This is known as vandalism. We extend the framework presented in (Potthast, Stein, and…

Information Retrieval · Computer Science 2012-10-23 Santiago M. Mola-Velasco

We introduce a next-generation vandalism detection system for Wikidata, one of the largest open-source structured knowledge bases on the Web. Wikidata is highly complex: its items incorporate an ever-expanding universe of factual triples…

Computation and Language · Computer Science 2025-05-26 Mykola Trokhymovych , Lydia Pintscher , Ricardo Baeza-Yates , Diego Saez-Trumper

A bag-of-words based probabilistic classifier is trained using regularized logistic regression to detect vandalism in the English Wikipedia. Isotonic regression is used to calibrate the class membership probabilities. Learning curve,…

Machine Learning · Computer Science 2010-01-06 Amit Belani

Vision-based inspection algorithms have significantly contributed to quality control in industrial settings, particularly in addressing structural defects like dent and contamination which are prevalent in mass production. Extensive…

Computer Vision and Pattern Recognition · Computer Science 2024-07-26 Kangil Lee , Geonuk Kim

This paper presents a novel design of the system aimed at supporting the Wikipedia community in addressing vandalism on the platform. To achieve this, we collected a massive dataset of 47 languages, and applied advanced filtering and…

Machine Learning · Computer Science 2023-06-05 Mykola Trokhymovych , Muniza Aslam , Ai-Jou Chou , Ricardo Baeza-Yates , Diego Saez-Trumper

Wikipedia is the largest online encyclopedia that allows anyone to edit articles. In this paper, we propose the use of deep learning to detect vandals based on their edit history. In particular, we develop a multi-source long-short term…

Cryptography and Security · Computer Science 2017-06-06 Shuhan Yuan , Panpan Zheng , Xintao Wu , Yang Xiang

With the continuous increase of data daily published in knowledge bases across the Web, one of the main issues is regarding information relevance. In most knowledge bases, a triple (i.e., a statement composed by subject, predicate, and…

Information Retrieval · Computer Science 2017-12-25 Edgard Marx , Tommaso Soru , André Valdestilhas

OpenStreetMap is a unique source of openly available worldwide map data, increasingly adopted in real-world applications. Vandalism detection in OpenStreetMap is critical and remarkably challenging due to the large scale of the dataset, the…

Machine Learning · Computer Science 2022-03-22 Nicolas Tempelmeier , Elena Demidova

Research on vandalism in Wikipedia has been of interest for the last decade. This paper performs a literature review on the subject, with the goal of identifying the main research topics and approaches, methods and techniques used. 67…

Digital Libraries · Computer Science 2016-06-20 Jesús Tramullas , Piedad Garrido-Picazo , Ana I. Sánchez-Casabón

Online fake news profoundly distorts public judgment and erodes trust in social platforms. While existing detectors achieve competitive performance on benchmark datasets, they remain notably vulnerable to malicious comments designed…

Machine Learning · Computer Science 2026-02-06 Zhao Tong , Chunlin Gong , Yimeng Gu , Haichao Shi , Qiang Liu , Shu Wu , Xiao-Yu Zhang

Wikidata is currently the largest open knowledge graph on the web, encompassing over 120 million entities. It integrates data from various domain-specific databases and imports a substantial amount of content from Wikipedia, while also…

Computation and Language · Computer Science 2026-01-06 Shixiong Zhao , Hideaki Takeda

LLM watermarking has attracted attention as a promising way to detect AI-generated content, with some works suggesting that current schemes may already be fit for deployment. In this work we dispute this claim, identifying watermark…

Machine Learning · Computer Science 2024-06-25 Nikola Jovanović , Robin Staab , Martin Vechev

Detecting controversy in general web pages is a daunting task, but increasingly essential to efficiently moderate discussions and effectively filter problematic content. Unfortunately, controversies occur across many topics and domains,…

Information Retrieval · Computer Science 2018-12-04 Jasper Linmans , Bob van de Velde , Evangelos Kanoulas

Data crowdsourcing is a data acquisition process where groups of voluntary contributors feed platforms with highly relevant data ranging from news, comments, and media to knowledge and classifications. It typically processes user-generated…

‹ Prev 1 2 3 10 Next ›