English
Related papers

Related papers: StRE: Self Attentive Edit Quality Prediction in Wi…

200 papers

The rise of AI-generated content in popular information sources raises significant concerns about accountability, accuracy, and bias amplification. Beyond directly impacting consumers, the widespread presence of this content poses questions…

Computation and Language · Computer Science 2024-10-11 Creston Brooks , Samuel Eggert , Denis Peskoff

In this work, we propose an automatic evaluation and comparison of the browsing behavior of Wikipedia readers that can be applied to any language editions of Wikipedia. As an example, we focus on English, French, and Russian languages…

Social and Information Networks · Computer Science 2020-02-18 Volodymyr Miz , Joëlle Hanna , Nicolas Aspert , Benjamin Ricaud , Pierre Vandergheynst

Wikipedia entity pages are a valuable source of information for direct consumption and for knowledge-base construction, update and maintenance. Facts in these entity pages are typically supported by references. Recent studies show that as…

Information Retrieval · Computer Science 2017-03-31 Besnik Fetahu , Katja Markert , Avishek Anand

Textual knowledge bases such as Wikipedia require considerable effort to keep up to date and consistent. While automated writing assistants could potentially ease this burden, the problem of suggesting edits grounded in external knowledge…

Computation and Language · Computer Science 2022-07-14 Robert L. Logan , Alexandre Passos , Sameer Singh , Ming-Wei Chang

The limited size of existing query-focused summarization datasets renders training data-driven summarization models challenging. Meanwhile, the manual construction of a query-focused summarization corpus is costly and time-consuming. In…

Computation and Language · Computer Science 2022-07-25 Haichao Zhu , Li Dong , Furu Wei , Bing Qin , Ting Liu

We evaluate the performance of transformer encoders with various decoders for information organization through a new task: generation of section headings for Wikipedia articles. Our analysis shows that decoders containing attention…

Computation and Language · Computer Science 2020-05-25 Anjalie Field , Sascha Rothe , Simon Baumgartner , Cong Yu , Abe Ittycheriah

Scene text editing (STE), which converts a text in a scene image into the desired text while preserving an original style, is a challenging task due to a complex intervention between text and style. In this paper, we propose a novel STE…

Computer Vision and Pattern Recognition · Computer Science 2022-05-03 Junyeop Lee , Yoonsik Kim , Seonghyeon Kim , Moonbin Yim , Seung Shin , Gayoung Lee , Sungrae Park

Given Wikipedia's role as a trusted source of high-quality, reliable content, concerns are growing about the proliferation of low-quality machine-generated text (MGT) produced by large language models (LLMs) on its platform. Reliable…

Computation and Language · Computer Science 2025-07-08 Gerrit Quaremba , Elizabeth Black , Denny Vrandečić , Elena Simperl

The use of domain knowledge is generally found to improve query efficiency in content filtering applications. In particular, tangible benefits have been achieved when using knowledge-based approaches within more specialized fields, such as…

Information Retrieval · Computer Science 2015-03-17 Pekka Malo , Pyry Siitari , Oskar Ahlgren , Jyrki Wallenius , Pekka Korhonen

We present a new, efficient method for automatically detecting severe conflicts `edit wars' in Wikipedia and evaluate this method on six different language WPs. We discuss how the number of edits, reverts, the length of discussions, the…

Machine Learning · Statistics 2023-01-05 Róbert Sumi , Taha Yasseri , András Rung , András Kornai , János Kertész

Wikipedia is one of the richest knowledge sources on the Web today. In order to facilitate navigating, searching, and maintaining its content, Wikipedia's guidelines state that all articles should be annotated with a so-called short…

Computation and Language · Computer Science 2023-02-20 Marija Sakota , Maxime Peyrard , Robert West

A model for the probabilistic function followed in Wikipedia edition is presented and compared with simulations and real data. It is argued that the probability to edit is proportional to the editor's number of previous editions…

Physics and Society · Physics 2021-10-27 Y. Gandica , F. Sampaio dos Aidos , J. Carvalho

Wikipedia articles (content pages) are commonly used corpora in Natural Language Processing (NLP) research, especially in low-resource languages other than English. Yet, a few research studies have studied the three Arabic Wikipedia…

Computation and Language · Computer Science 2024-04-02 Saied Alshahrani , Hesham Haroon , Ali Elfilali , Mariama Njie , Jeanna Matthews

The search for relevant information can be very frustrating for users who, unintentionally, use too general or inappropriate keywords to express their requests. To overcome this situation, query expansion techniques aim at transforming the…

Information Retrieval · Computer Science 2016-05-13 Joan Guisado-Gámez , Arnau Prat-Pérez , Josep Lluís Larriba-Pey

Experimenting with a new dataset of 1.6M user comments from a Greek news portal and existing datasets of English Wikipedia comments, we show that an RNN outperforms the previous state of the art in moderation. A deep,…

Computation and Language · Computer Science 2017-07-18 John Pavlopoulos , Prodromos Malakasiotis , Ion Androutsopoulos

Wikipedia has high-quality articles on a variety of topics and has been used in diverse research areas. In this study, a method is presented for using Wikipedia's editor information to build recommender systems in various domains that…

Information Retrieval · Computer Science 2023-06-16 Katsuhiko Hayashi

Explicit Semantic Analysis (ESA) is a technique used to represent a piece of text as a vector in the space of concepts, such as Articles found in Wikipedia. We propose a methodology to incorporate knowledge of Inter-relatedness between…

Computation and Language · Computer Science 2020-12-02 Naveen Elango , Pawan Prasad K

Content moderation in online platforms is crucial for ensuring activity therein adheres to existing policies, especially as these platforms grow. NLP research in this area has typically focused on automating some part of it given that it is…

Computation and Language · Computer Science 2024-08-13 Hsuvas Borkakoty , Luis Espinosa-Anke

Automated content moderation for collaborative knowledge hubs like Wikipedia or Wikidata is an important yet challenging task due to multiple factors. In this paper, we construct a database of discussions happening around articles marked…

Computation and Language · Computer Science 2025-03-14 Hsuvas Borkakoty , Luis Espinosa-Anke

The advent of large pre-trained language models has made it possible to make high-quality predictions on how to add or change a sentence in a document. However, the high branching factor inherent to text generation impedes the ability of…

Computation and Language · Computer Science 2021-06-15 Zeqiu Wu , Michel Galley , Chris Brockett , Yizhe Zhang , Bill Dolan