中文
相关论文

相关论文: Fair multilingual vandalism detection system for W…

200 篇论文

In this paper, we present the first results of our ongoing early-stage research on a realtime disaster detection and monitoring tool. Based on Wikipedia, it is language-agnostic and leverages user-generated multimedia content shared on…

社会与信息网络 · 计算机科学 2015-01-27 Thomas Steiner , Ruben Verborgh

We report on the Wikidata vandalism detection task at the WSDM Cup 2017. The task received five submissions for which this paper describes their evaluation and a comparison to state of the art baselines. Unlike previous work, we recast…

信息检索 · 计算机科学 2017-12-19 Stefan Heindorf , Martin Potthast , Gregor Engels , Benno Stein

Wikipedia is the largest online encyclopedia that allows anyone to edit articles. In this paper, we propose the use of deep learning to detect vandals based on their edit history. In particular, we develop a multi-source long-short term…

密码学与安全 · 计算机科学 2017-06-06 Shuhan Yuan , Panpan Zheng , Xintao Wu , Yang Xiang

In this paper, we focus on normative systems for online communities. The paper addresses the issue that arises when different community members interpret these norms in different ways, possibly leading to unexpected behavior in…

社会与信息网络 · 计算机科学 2024-02-08 Thiago Freitas dos Santos , Nardine Osman , Marco Schorlemmer

Distributed word representations, or word vectors, have recently been applied to many tasks in natural language processing, leading to state-of-the-art performance. A key ingredient to the successful application of these representations is…

计算与语言 · 计算机科学 2018-03-30 Edouard Grave , Piotr Bojanowski , Prakhar Gupta , Armand Joulin , Tomas Mikolov

We propose a simple, yet effective, approach towards inducing multilingual taxonomies from Wikipedia. Given an English taxonomy, our approach leverages the interlanguage links of Wikipedia followed by character-level classifiers to induce…

计算与语言 · 计算机科学 2017-09-13 Amit Gupta , Rémi Lebret , Hamza Harkous , Karl Aberer

Lack of diverse perspectives causes neutrality bias in Wikipedia content leading to millions of worldwide readers getting exposed by potentially inaccurate information. Hence, neutrality bias detection and mitigation is a critical problem.…

计算与语言 · 计算机科学 2023-12-27 Ankita Maity , Anubhav Sharma , Rudra Dhar , Tushar Abhishek , Manish Gupta , Vasudeva Varma

The different Wikipedia language editions vary dramatically in how comprehensive they are. As a result, most language editions contain only a small fraction of the sum of information that exists across all Wikipedias. In this paper, we…

社会与信息网络 · 计算机科学 2016-04-13 Ellery Wulczyn , Robert West , Leila Zia , Jure Leskovec

Wikipedia's perceived high quality and broad language coverage have established it as a fundamental resource in NLP. However, in recent years, such assumptions of high quality have become the subject of scrutiny in low-resource and…

Several studies have used Wikipedia (WP) data-set to analyse worldwide human preferences by languages. However, those studies could suffer from bias related to exceptional social circumstances. Any massive event promoting the exceptional…

物理与社会 · 物理学 2022-05-17 Julien Assuied , Yérali Gandica

Machine learning systems are ubiquitous in various kinds of digital applications and have a huge impact on our everyday life. But a lack of explainability and interpretability of such systems hinders meaningful participation by people,…

We present a new, efficient method for automatically detecting severe conflicts `edit wars' in Wikipedia and evaluate this method on six different language WPs. We discuss how the number of edits, reverts, the length of discussions, the…

机器学习 · 统计学 2023-01-05 Róbert Sumi , Taha Yasseri , András Rung , András Kornai , János Kertész

We release a corpus of 43 million atomic edits across 8 languages. These edits are mined from Wikipedia edit history and consist of instances in which a human editor has inserted a single contiguous phrase into, or deleted a single…

计算与语言 · 计算机科学 2018-08-29 Manaal Faruqui , Ellie Pavlick , Ian Tenney , Dipanjan Das

Multilingualism is common offline, but we have a more limited understanding of the ways multilingualism is displayed online and the roles that multilinguals play in the spread of content between speakers of different languages. We take a…

社会与信息网络 · 计算机科学 2016-06-14 Suin Kim , Sungjoon Park , Scott A. Hale , Sooyoung Kim , Jeongmin Byun , Alice Oh

Hyperlinks constitute the backbone of the Web; they enable user navigation, information discovery, content ranking, and many other crucial services on the Internet. In particular, hyperlinks found within Wikipedia allow the readers to…

计算机与社会 · 计算机科学 2021-06-01 Martin Gerlach , Marshall Miller , Rita Ho , Kosta Harlan , Djellel Difallah

Wikipedia (WP) as a collaborative, dynamical system of humans is an appropriate subject of social studies. Each single action of the members of this society, i.e. editors, is well recorded and accessible. Using the cumulative data of 34…

物理与社会 · 物理学 2023-01-05 Taha Yasseri , Róbert Sumi , János Kertész

Wikidata is the new, large-scale knowledge base of the Wikimedia Foundation. As it can be edited by anyone, entries frequently get vandalized, leading to the possibility that it might spread of falsified information if such posts are not…

信息检索 · 计算机科学 2017-12-20 Qi Zhu , Hongwei Ng , Liyuan Liu , Ziwei Ji , Bingjie Jiang , Jiaming Shen , Huan Gui

Wikipedia is the largest web repository of free knowledge. Volunteer editors devote time and effort to creating and expanding articles in more than 300 language editions. As content quality varies from article to article, editors also spend…

计算机与社会 · 计算机科学 2024-04-16 Paramita Das , Isaac Johnson , Diego Saez-Trumper , Pablo Aragón

The WSDM Cup 2017 was a data mining challenge held in conjunction with the 10th International Conference on Web Search and Data Mining (WSDM). It addressed key challenges of knowledge bases today: quality assurance and entity search. For…

信息检索 · 计算机科学 2017-12-29 Martin Potthast , Stefan Heindorf , Hannah Bast

We introduce a new reading comprehension dataset, dubbed MultiWikiQA, which covers 306 languages and has 1,220,757 samples in total. We start with Wikipedia articles, which also provide the context for the dataset samples, and use an LLM to…

计算与语言 · 计算机科学 2026-03-05 Dan Saattrup Smart