English
Related papers

Related papers: A practical approach to language complexity: a Wik…

200 papers

I started this work with the hope of generating a text synthesizer (like a musical synthesizer) that can imitate certain linguistic styles. Most of the report focuses on text simplification using statistical machine translation (SMT)…

Computation and Language · Computer Science 2017-03-28 Yohan Jo

We conducted a global comparative analysis of the coverage of American topics in different language versions of Wikipedia, using over 90 million Wikidata items and 40 million Wikipedia articles in 58 languages. Our study aimed to…

Information Retrieval · Computer Science 2023-07-28 Piotr Konieczny , Włodzimierz Lewoniewski

In today's world, we follow news which is distributed globally. Significant events are reported by different sources and in different languages. In this work, we address the problem of tracking of events in a large multilingual stream.…

Information Retrieval · Computer Science 2015-12-23 Jan Rupnik , Andrej Muhic , Gregor Leban , Primoz Skraba , Blaz Fortuna , Marko Grobelnik

This paper aims to review the fiercely discussed question of whether the ranking of Wikipedia articles in search engines is justified by the quality of the articles. After an overview of current research on information quality in Wikipedia,…

Information Retrieval · Computer Science 2011-09-06 Dirk Lewandowski , Ulrike Spree

The Internet-based encyclopaedia Wikipedia has grown to become one of the most visited web-sites on the Internet. However, critics have questioned the quality of entries, and an empirical study has shown Wikipedia to contain errors in a…

Digital Libraries · Computer Science 2011-01-04 Finn Aarup Nielsen

Although Wikipedia is the largest multilingual encyclopedia, it remains inherently incomplete. There is a significant disparity in the quality of content between high-resource languages (HRLs, e.g., English) and low-resource languages…

Computation and Language · Computer Science 2024-12-10 Paramita Das , Amartya Roy , Ritabrata Chakraborty , Animesh Mukherjee

Wikipedia is the largest open knowledge corpus, widely used worldwide and serving as a key resource for training large language models (LLMs) and retrieval-augmented generation (RAG) systems. Ensuring its accuracy is therefore critical. But…

Computation and Language · Computer Science 2025-09-30 Sina J. Semnani , Jirayu Burapacheep , Arpandeep Khatua , Thanawan Atchariyachanvanit , Zheng Wang , Monica S. Lam

We revisit the phenomenon of syntactic complexity convergence in conversational interaction, originally found for English dialogue, which has theoretical implication for dialogical concepts such as mutual understanding. We use a modified…

Computation and Language · Computer Science 2024-08-23 Yu Wang , Hendrik Buschmeier

Wikipedia is the largest source of free encyclopedic knowledge and one of the most visited sites on the Web. To increase reader understanding of the article, Wikipedia editors add images within the text of the article's body. However,…

Computers and Society · Computer Science 2021-12-06 Daniele Rama , Tiziano Piccardi , Miriam Redi , Rossano Schifanella

We present a language complexity analysis of World of Warcraft (WoW) community texts, which we compare to texts from a general corpus of web English. Results from several complexity types are presented, including lexical diversity, density,…

Computation and Language · Computer Science 2015-02-11 Simon Šuster

Detecting controversy in general web pages is a daunting task, but increasingly essential to efficiently moderate discussions and effectively filter problematic content. Unfortunately, controversies occur across many topics and domains,…

Information Retrieval · Computer Science 2018-12-04 Jasper Linmans , Bob van de Velde , Evangelos Kanoulas

Automatic quality evaluation of Web information is a task with many fields of applications and of great relevance, especially in critical domains like the medical one. We move from the intuition that the quality of content of medical Web…

Information Retrieval · Computer Science 2016-03-08 Vittoria Cozza , Marinella Petrocchi , Angelo Spognardi

A model for the probabilistic function followed in Wikipedia edition is presented and compared with simulations and real data. It is argued that the probability to edit is proportional to the editor's number of previous editions…

Physics and Society · Physics 2021-10-27 Y. Gandica , F. Sampaio dos Aidos , J. Carvalho

Wikipedia is one of the most popular sites on the Web, with millions of users relying on it to satisfy a broad range of information needs every day. Although it is crucial to understand what exactly these needs are in order to be able to…

Social and Information Networks · Computer Science 2017-03-17 Philipp Singer , Florian Lemmerich , Robert West , Leila Zia , Ellery Wulczyn , Markus Strohmaier , Jure Leskovec

Improving pretraining data quality and size is known to boost downstream performance, but the role of text complexity--how hard a text is to read--remains less explored. We reduce surface-level complexity (shorter sentences, simpler words,…

Computation and Language · Computer Science 2025-10-07 Dan John Velasco , Matthew Theodore Roque

The voluntary process of Wikipedia edition provides an environment where the outcome is clearly a collective product of interactions involving a large number of people. We propose a simple agent-based model, developed from real data, to…

Physics and Society · Physics 2015-06-22 Y. Gandica , F. Sampaio dos Aidos , J. Carvalho

Wikipedia is one of the most visited websites globally, yet its role beyond its own platform remains largely unexplored. In this paper, we present the first large-scale analysis of how Wikipedia is referenced across the Web. Using a dataset…

Social and Information Networks · Computer Science 2025-05-23 Veniamin Veselovsky , Tiziano Piccardi , Ashton Anderson , Robert West , Akhil Arora

We introduce a new reading comprehension dataset, dubbed MultiWikiQA, which covers 306 languages and has 1,220,757 samples in total. We start with Wikipedia articles, which also provide the context for the dataset samples, and use an LLM to…

Computation and Language · Computer Science 2026-03-05 Dan Saattrup Smart

Wikipedia is a goldmine of information; not just for its many readers, but also for the growing community of researchers who recognize it as a resource of exceptional scale and utility. It represents a vast investment of manual effort and…

Artificial Intelligence · Computer Science 2009-05-10 Olena Medelyan , David Milne , Catherine Legg , Ian H. Witten

While Wikipedia exists in 287 languages, its content is unevenly distributed among them. In this work, we investigate the generation of open domain Wikipedia summaries in underserved languages using structured data from Wikidata. To this…

Computation and Language · Computer Science 2018-05-01 Lucie-Aimée Kaffee , Hady Elsahar , Pavlos Vougiouklis , Christophe Gravier , Frédérique Laforest , Jonathon Hare , Elena Simperl