English
Related papers

Related papers: Measuring Wikipedia Article Quality in One Dimensi…

200 papers

Sections are the building blocks of Wikipedia articles. They enhance readability and can be used as a structured entry point for creating and expanding articles. Structuring a new or already existing Wikipedia article with sections is a…

Information Retrieval · Computer Science 2018-05-07 Tiziano Piccardi , Michele Catasta , Leila Zia , Robert West

Social biases on Wikipedia, a widely-read global platform, could greatly influence public opinion. While prior research has examined man/woman gender bias in biography articles, possible influences of other demographic attributes limit…

Computation and Language · Computer Science 2022-02-10 Anjalie Field , Chan Young Park , Kevin Z. Lin , Yulia Tsvetkov

Classifier-based Quality Filtering has recently emerged as a fundamental technique in constructing pre-training corpora. The ability to deploy a single model that can replace or supplement a set of heuristics has proven effective across…

Computation and Language · Computer Science 2026-05-25 Mateusz Klimaszewski , Piotr Andruszkiewicz

In this work, we compare two simple methods of tagging scientific publications with labels reflecting their content. As a first source of labels Wikipedia is employed, second label set is constructed from the noun phrases occurring in the…

Computation and Language · Computer Science 2014-11-04 Michał Łopuszyński , Łukasz Bolikowski

We review some recent endeavors and add some new results to characterize and understand underlying mechanisms in Wikipedia (WP), the paradigmatic example of collaborative value production. We analyzed the statistics of editorial activity in…

Physics and Society · Physics 2023-01-05 Taha Yasseri , János Kertész

Experimental design techniques such as active search and Bayesian optimization are widely used in the natural sciences for data collection and discovery. However, existing techniques tend to favor exploitation over exploration of the search…

Machine Learning · Statistics 2024-05-07 Quan Nguyen , Adji Bousso Dieng

Wikipedia is a huge opportunity for machine learning, being the largest semi-structured base of knowledge available. Because of this, many works examine its contents, and focus on structuring it in order to make it usable in learning tasks,…

Machine Learning · Computer Science 2020-01-23 Tiphaine Viard , Thomas McLachlan , Hamidreza Ghader , Satoshi Sekine

Ordinal classification problems, where labels exhibit a natural order, are prevalent in high-stakes fields such as medicine and finance. Accurate uncertainty quantification, including the decomposition into aleatoric (inherent variability)…

Machine Learning · Computer Science 2025-07-02 Stefan Haas , Eyke Hüllermeier

Eigenfactor.org, a journal evaluation tool which uses an iterative algorithm to weight citations (similar to the PageRank algorithm used for Google) has been proposed as a more valid method for calculating the impact of journals. The…

Digital Libraries · Computer Science 2009-09-29 Philip M. Davis

The use of domain knowledge is generally found to improve query efficiency in content filtering applications. In particular, tangible benefits have been achieved when using knowledge-based approaches within more specialized fields, such as…

Information Retrieval · Computer Science 2015-03-17 Pekka Malo , Pyry Siitari , Oskar Ahlgren , Jyrki Wallenius , Pekka Korhonen

Wikipedia is the largest open knowledge corpus, widely used worldwide and serving as a key resource for training large language models (LLMs) and retrieval-augmented generation (RAG) systems. Ensuring its accuracy is therefore critical. But…

Computation and Language · Computer Science 2025-09-30 Sina J. Semnani , Jirayu Burapacheep , Arpandeep Khatua , Thanawan Atchariyachanvanit , Zheng Wang , Monica S. Lam

Information presented in Wikipedia articles must be attributable to reliable published sources in the form of references. This study examines over 5 million Wikipedia articles to assess the reliability of references in multiple language…

Computers and Society · Computer Science 2023-09-06 Aitolkyn Baigutanova , Diego Saez-Trumper , Miriam Redi , Meeyoung Cha , Pablo Aragón

Nowadays, editors tend to separate different subtopics of a long Wiki-pedia article into multiple sub-articles. This separation seeks to improve human readability. However, it also has a deleterious effect on many Wikipedia-based tasks that…

Information Retrieval · Computer Science 2019-06-24 Muhao Chen , Changping Meng , Gang Huang , Carlo Zaniolo

With the spread of online social networks, it is more and more difficult to monitor all the user-generated content. Automating the moderation process of the inappropriate exchange content on Internet has thus become a priority task. Methods…

Computation and Language · Computer Science 2021-01-19 Noé Cecillon , Vincent Labatut , Richard Dufour , Georges Linares

Purpose: Scholars often aim to conduct high quality research and their success is judged primarily by peer reviewers. Research quality is difficult for either group to identify, however, and misunderstandings can reduce the efficiency of…

Digital Libraries · Computer Science 2022-12-13 Mike Thelwall , Kayvan Kousha , Mahshid Abdoli , Emma Stuart , Meiko Makita , Paul Wilson , Jonathan Levitt

Wikipedia is an invaluable resource for factual information about a wide range of entities. However, the quality of articles on less-known entities often lags behind that of the well-known ones. This study proposes a novel approach to…

Computation and Language · Computer Science 2025-02-18 Sayantan Adak , Pauras Mangesh Meher , Paramita Das , Animesh Mukherjee

Although there is an emerging trend towards generating embeddings for primarily unstructured data and, recently, for structured data, no systematic suite for measuring the quality of embeddings has been proposed yet. This deficiency is…

Computation and Language · Computer Science 2020-05-11 Faisal Alshargi , Saeedeh Shekarpour , Tommaso Soru , Amit Sheth

Peer production projects such as Wikipedia or open-source software development allow volunteers to collectively create knowledge based products. The inclusive nature of such projects poses difficult challenges for ensuring trustworthiness…

Computer Science and Game Theory · Computer Science 2014-02-05 S. Anand , Ofer Arazy , Narayan Mandayam , Oded Nov

We study text reuse related to Wikipedia at scale by compiling the first corpus of text reuse cases within Wikipedia as well as without (i.e., reuse of Wikipedia text in a sample of the Common Crawl). To discover reuse beyond verbatim copy…

Information Retrieval · Computer Science 2018-12-24 Milad Alshomary , Michael Völske , Tristan Licht , Henning Wachsmuth , Benno Stein , Matthias Hagen , Martin Potthast

A sequence of recent papers has considered the role of measurement scales in information retrieval (IR) experimentation, and presented the argument that (only) uniform-step interval scales should be used, and hence that well-known metrics…

Information Retrieval · Computer Science 2022-07-08 Alistair Moffat