English
Related papers

Related papers: Automatically Labeling Low Quality Content on Wiki…

200 papers

Wikipedia plays a crucial role in the integrity of the Web. This work analyzes the reliability of this global encyclopedia through the lens of its references. We operationalize the notion of reference quality by defining reference need…

Digital Libraries · Computer Science 2023-03-10 Aitolkyn Baigutanova , Jaehyeon Myung , Diego Saez-Trumper , Ai-Jou Chou , Miriam Redi , Changwook Jung , Meeyoung Cha

In this paper, we present a comprehensive analysis and monitoring framework for the impact of Large Language Models (LLMs) on Wikipedia, examining the evolution of Wikipedia through existing data and using simulations to explore potential…

Computation and Language · Computer Science 2026-03-03 Siming Huang , Yuliang Xu , Mingmeng Geng , Yao Wan , Dongping Chen

In this paper we discuss several issues related to automated text classification of web sites. We analyze the nature of web content and metadata in relation to requirements for text features. We find that HTML metatags are a good source of…

Information Retrieval · Computer Science 2007-05-23 John M. Pierre

Verifiability is one of the core editing principles in Wikipedia, where editors are encouraged to provide citations for the added statements. Statements can be any arbitrary piece of text, ranging from a sentence up to a paragraph. However,…

Computation and Language · Computer Science 2018-05-01 Besnik Fetahu

Encyclopedic queries express the intent of obtaining information typically available in encyclopedias, such as biographical, geographical or historical facts. In this paper, we train a classifier for detecting the encyclopedic intent of web…

Information Retrieval · Computer Science 2015-12-01 Pedro Saleiro , Luís Sarmento

Semantic annotations have to satisfy quality constraints to be useful for digital libraries, which is particularly challenging on large and diverse datasets. Confidence scores of multi-label classification methods typically refer only to…

Information Retrieval · Computer Science 2018-06-08 Martin Toepfer , Christin Seifert

An important editing policy in Wikipedia is to provide citations for added statements in Wikipedia pages, where statements can be arbitrary pieces of text, ranging from a sentence to a paragraph. In many cases citations are either outdated…

Information Retrieval · Computer Science 2017-04-26 Besnik Fetahu , Katja Markert , Wolfgang Nejdl , Avishek Anand

AI tools are increasingly deployed in community contexts. However, datasets used to evaluate AI are typically created by developers and annotators outside a given community, which can yield misleading conclusions about AI performance. How…

Human-Computer Interaction · Computer Science 2024-02-23 Tzu-Sheng Kuo , Aaron Halfaker , Zirui Cheng , Jiwoo Kim , Meng-Hsin Wu , Tongshuang Wu , Kenneth Holstein , Haiyi Zhu

This paper presents an approach to classify documents in any language into an English topical label space, without any text categorization training data. The approach, Cross-Lingual Dataless Document Classification (CLDDC) relies on mapping…

Computation and Language · Computer Science 2016-11-15 Yangqiu Song , Stephen Mayhew , Dan Roth

Wikipedia is among the largest examples of collective intelligence on the Web with over 61 million articles covering over 320 languages. Although edited and maintained by an active workforce of human volunteers, Wikipedia is highly reliant…

Human-Computer Interaction · Computer Science 2025-09-29 Neal Reeves , Elena Simperl

Topics generated by topic models are typically represented as list of terms. To reduce the cognitive overhead of interpreting these topics for end-users, we propose labelling a topic with a succinct phrase that summarises its theme or idea.…

Computation and Language · Computer Science 2016-12-26 Shraey Bhatia , Jey Han Lau , Timothy Baldwin

Over the last few years, verifying the credibility of information sources has become a fundamental need to combat disinformation. Here, we present a language-agnostic model designed to assess the reliability of web domains as sources in…

Social and Information Networks · Computer Science 2025-11-21 Jacopo D'Ignazi , Andreas Kaltenbrunner , Yelena Mejova , Michele Tizzani , Kyriaki Kalimeri , Mariano Beiró , Pablo Aragón

In this paper we address the challenge of assessing the quality of Wikipedia pages using scores derived from edit contribution and contributor authoritativeness measures. The hypothesis is that pages with significant contributions from…

Social and Information Networks · Computer Science 2013-10-25 Xiangju Qin , Pádraig Cunningham

The vast amount of online information today poses challenges for non-English speakers, as much of it is concentrated in high-resource languages such as English and French. Wikipedia reflects this imbalance, with content in low-resource…

Computation and Language · Computer Science 2025-04-08 Siddharth Khincha , Tushar Kataria , Ankita Anand , Dan Roth , Vivek Gupta

The rise of AI-generated content in popular information sources raises significant concerns about accountability, accuracy, and bias amplification. Beyond directly impacting consumers, the widespread presence of this content poses questions…

Computation and Language · Computer Science 2024-10-11 Creston Brooks , Samuel Eggert , Denis Peskoff

Wikipedia is a useful knowledge source that benefits many applications in language processing and knowledge representation. An important feature of Wikipedia is that of categories. Wikipedia pages are assigned different categories according…

Computation and Language · Computer Science 2017-04-26 Yanqing Chen , Steven Skiena

Most existing large language models (LLMs) are expensive to adapt after deployment, especially when a task requires newly produced information or niche domain knowledge. Recent work has shown that, by manipulating and optimizing their…

Computation and Language · Computer Science 2026-05-15 Zeyu Huang , Adhiguna Kuncoro , Qixuan Feng , Jiajun Shen , Lucio Dery , Arthur Szlam , Marc'Aurelio Ranzato

Auto-annotation by ensemble of models is an efficient method of learning on unlabeled data. Wrong or inaccurate annotations generated by the ensemble may lead to performance degradation of the trained model. To deal with this problem we…

Computer Vision and Pattern Recognition · Computer Science 2024-03-14 Dror Simon , Miriam Farber , Roman Goldenberg

Most of the existing information retrieval systems are based on bag of words model and are not equipped with common world knowledge. Work has been done towards improving the efficiency of such systems by using intelligent algorithms to…

Artificial Intelligence · Computer Science 2015-03-17 Pekka Malo , Pyry Siitari , Ankur Sinha

Sections are the building blocks of Wikipedia articles. They enhance readability and can be used as a structured entry point for creating and expanding articles. Structuring a new or already existing Wikipedia article with sections is a…

Information Retrieval · Computer Science 2018-05-07 Tiziano Piccardi , Michele Catasta , Leila Zia , Robert West