English
Related papers

Related papers: MediaWiki Grammar Recovery

200 papers

Wikidata has been increasingly adopted by many communities for a wide variety of applications, which demand high-quality knowledge to deliver successful results. In this paper, we develop a framework to detect and analyze low-quality…

Artificial Intelligence · Computer Science 2021-11-22 Kartik Shenoy , Filip Ilievski , Daniel Garijo , Daniel Schwabe , Pedro Szekely

Wikipedia is the largest web repository of free knowledge. Volunteer editors devote time and effort to creating and expanding articles in more than 300 language editions. As content quality varies from article to article, editors also spend…

Computers and Society · Computer Science 2024-04-16 Paramita Das , Isaac Johnson , Diego Saez-Trumper , Pablo Aragón

Data leakage has been identified in 648 published machine learning papers across 30 scientific fields. The knowledge to prevent it exists; the tools do not enforce it. This paper presents a grammar - eight typed primitives, a directed…

Machine Learning · Computer Science 2026-04-07 Simon Roth

Wikipedia is the world's largest online encyclopedia, but maintaining article quality through collaboration is challenging. Wikipedia designed a quality scale, but with such a manual assessment process, many articles remain unassessed. We…

Computation and Language · Computer Science 2023-10-04 Pedro Miguel Moás , Carla Teixeira Lopes

Fake information poses one of the major threats for society in the 21st century. Identifying misinformation has become a key challenge due to the amount of fake news that is published daily. Yet, no approach is established that addresses…

Information Retrieval · Computer Science 2021-03-30 Vishwani Gupta , Katharina Beckh , Sven Giesselbach , Dennis Wegener , Tim Wirtz

Wikipedia serves as a key infrastructure for public access to scientific knowledge, but it faces challenges in maintaining the credibility of cited sources--especially when scientific papers are retracted. This paper investigates how…

Human-Computer Interaction · Computer Science 2026-01-27 Haohan Shi , Yulin Yu , Daniel M. Romero , Emőke-Ágnes Horvát

Grammar refers to the system of rules that governs the structural organization and the semantic relations among linguistic units such as sentences, phrases, and words within a given language. In natural language processing, there remains a…

Computation and Language · Computer Science 2026-02-24 Lujun Li , Yewei Song , Lama Sleem , Yiqun Wang , Yangjie Xu , Cedric Lothritz , Niccolo Gentile , Radu State , Tegawende F. Bissyande , Jacques Klein

Grammar convergence is a method that helps discovering relationships between different grammars of the same language or different language versions. The key element of the method is the operational, transformation-based representation of…

Programming Languages · Computer Science 2011-07-20 Ralf Lämmel , Vadim Zaytsev

Wikipedia articles aim to be definitive sources of encyclopedic content. Yet, only 0.6% of Wikipedia articles have high quality according to its quality scale due to insufficient number of Wikipedia editors and enormous number of articles.…

Social and Information Networks · Computer Science 2021-08-06 Sumit Asthana , Sabrina Tobar Thommel , Aaron Lee Halfaker , Nikola Banovic

Wikipedia's perceived high quality and broad language coverage have established it as a fundamental resource in NLP. However, in recent years, such assumptions of high quality have become the subject of scrutiny in low-resource and…

We have implemented the Memento MediaWiki Extension Version 2.0, which brings the Memento Protocol to MediaWiki, used by Wikipedia and the Wikimedia Foundation. Test results show that the extension has a negligible impact on performance.…

Digital Libraries · Computer Science 2014-06-17 Shawn M. Jones , Michael L. Nelson , Harihar Shankar , Herbert Van de Sompel

Large language models demonstrate limited capability in proficiency-controlled sentence simplification, particularly when simplifying across large readability levels. We propose a framework that decomposes complex simplifications into…

Computation and Language · Computer Science 2026-02-10 Jingshen Zhang , Xin Ying Qiu , Lifang Lu , Zhuhua Huang , Yutao Hu , Yuechang Wu , JunYu Lu

Idioms are common in everyday language, but often pose a challenge to translators because their meanings do not follow from the meanings of their parts. Despite significant advances, machine translation systems still struggle to translate…

Computation and Language · Computer Science 2023-10-24 Emmy Liu , Aditi Chaudhary , Graham Neubig

Semantic communication has emerged as the breakthrough beyond the Shannon theorem by transmitting and receiving semantic information instead of data bits or symbols regardless of its content. This paper proposes a two-stage reconstruction…

Information Theory · Computer Science 2022-09-13 Trinh Van Chien , Le Hong Phong , Dao Xuan Phuc , Nguyen Tien Hoa

Queries to large language models (LLMs) can be divided into two parts: the instruction/question and the accompanying context. The context for retrieval-augmented generation (RAG) systems in most benchmarks comes from Wikipedia-like texts…

Computation and Language · Computer Science 2025-07-01 Benjamin Reichman , Adar Avsian , Kartik Talamadupula , Toshish Jawale , Larry Heck

The users of endangered languages struggle to thrive in a digitally-mediated world. We have developed an automated method for assessing how well every language recognized by ISO 639 is faring in terms of digital language support. The…

Computation and Language · Computer Science 2022-09-28 Gary F. Simons , Abbey L. Thomas , Chad K. White

Classifier-based Quality Filtering has recently emerged as a fundamental technique in constructing pre-training corpora. The ability to deploy a single model that can replace or supplement a set of heuristics has proven effective across…

Computation and Language · Computer Science 2026-05-25 Mateusz Klimaszewski , Piotr Andruszkiewicz

Grammatical error classification plays a crucial role in language learning systems, but existing classification taxonomies often lack rigorous validation, leading to inconsistencies and unreliable feedback. In this paper, we revisit…

Computation and Language · Computer Science 2025-02-19 Deqing Zou , Jingheng Ye , Yulu Liu , Yu Wu , Zishan Xu , Yinghui Li , Hai-Tao Zheng , Bingxu An , Zhao Wei , Yong Xu

We introduce a new reading comprehension dataset, dubbed MultiWikiQA, which covers 306 languages and has 1,220,757 samples in total. We start with Wikipedia articles, which also provide the context for the dataset samples, and use an LLM to…

Computation and Language · Computer Science 2026-03-05 Dan Saattrup Smart

Webpages have been a rich, scalable resource for vision-language and language only tasks. Yet only pieces of webpages are kept in existing datasets: image-caption pairs, long text articles, or raw HTML, never all in one place. Webpage tasks…

Computation and Language · Computer Science 2023-10-23 Andrea Burns , Krishna Srinivasan , Joshua Ainslie , Geoff Brown , Bryan A. Plummer , Kate Saenko , Jianmo Ni , Mandy Guo