English
Related papers

Related papers: Language-agnostic Topic Classification for Wikiped…

200 papers

A simple dynamical model of collective edit activity of Wikipedia articles and their content evolution is introduced. Based on the recent empirical findings, each editor in the model is characterized by an ability to make content edit,…

Physics and Society · Physics 2023-04-25 Takashi Shimada , Fumiko Ogushi , Janos Torok , Janos Kertesz , Kimmo Kaski

With the development of deep learning and natural language processing techniques, pre-trained language models have been widely used to solve information retrieval (IR) problems. Benefiting from the pre-training and fine-tuning paradigm,…

Information Retrieval · Computer Science 2024-01-02 Weihang Su , Qingyao Ai , Xiangsheng Li , Jia Chen , Yiqun Liu , Xiaolong Wu , Shengluan Hou

Although Wikipedia is the largest multilingual encyclopedia, it remains inherently incomplete. There is a significant disparity in the quality of content between high-resource languages (HRLs, e.g., English) and low-resource languages…

Computation and Language · Computer Science 2024-12-10 Paramita Das , Amartya Roy , Ritabrata Chakraborty , Animesh Mukherjee

We present a new concept - Wikiometrics - the derivation of metrics and indicators from Wikipedia. Wikipedia provides an accurate representation of the real world due to its size, structure, editing policy and popularity. We demonstrate an…

Digital Libraries · Computer Science 2016-01-11 Gilad Katz , Lior Rokach

Wikipedia is a well-known platform for disseminating knowledge, and scientific sources, such as journal articles, play a critical role in supporting its mission. The open access movement aims to make scientific knowledge openly available,…

Digital Libraries · Computer Science 2024-10-17 Puyu Yang , Ahad Shoaib , Robert West , Giovanni Colavizza

The Internet-based encyclopaedia Wikipedia has grown to become one of the most visited web-sites on the Internet. However, critics have questioned the quality of entries, and an empirical study has shown Wikipedia to contain errors in a…

Digital Libraries · Computer Science 2011-01-04 Finn Aarup Nielsen

To achieve equitable performance across languages, large language models (LLMs) must be able to abstract knowledge beyond the language in which it was learnt. However, the current literature lacks reliable ways to measure LLMs' capability…

The frequency of a web search keyword generally reflects the degree of public interest in a particular subject matter. Search logs are therefore useful resources for trend analysis. However, access to search logs is typically restricted to…

Social and Information Networks · Computer Science 2015-09-09 Mitsuo Yoshida , Yuki Arase , Takaaki Tsunoda , Mikio Yamamoto

Adequate representation of natural language semantics requires access to vast amounts of common sense and domain-specific world knowledge. Prior work in the field was based on purely statistical techniques that did not make use of…

Computation and Language · Computer Science 2014-01-23 Evgeniy Gabrilovich , Shaul Markovitch

Geopolitics focuses on political power in relation to geographic space. Interactions among world countries have been widely studied at various scales, observing economic exchanges, world history or international politics among others. This…

Social and Information Networks · Computer Science 2017-07-21 Klaus M. Frahm , Samer El Zant , Katia Jaffrès-Runser , Dima L. Shepelyansky

We propose a dynamic map of knowledge generated from Wikipedia pages and the Web URLs contained therein. GalaxySearch provides answers to the questions we don't know how to ask, by constructing a semantic network of the most relevant pages…

Social and Information Networks · Computer Science 2012-04-17 Hauke Fuehres , Peter A. Gloor , Michael Henninger , Reto Kleeb , Keiichi Nemoto

Wikipedia is a huge global repository of human knowledge, that can be leveraged to investigate interwinements between cultures. With this aim, we apply methods of Markov chains and Google matrix, for the analysis of the hyperlink networks…

Social and Information Networks · Computer Science 2015-03-06 Young-Ho Eom , Pablo Aragón , David Laniado , Andreas Kaltenbrunner , Sebastiano Vigna , Dima L. Shepelyansky

We present our work on aligning the Unified Medical Language System (UMLS) to Wikipedia, to facilitate manual alignment of the two resources. We propose a cross-lingual neural reranking model to match a UMLS concept with a Wikipedia page,…

Computation and Language · Computer Science 2020-11-03 Afshin Rahimi , Timothy Baldwin , Karin Verspoor

Since its inception six years ago, the online encyclopedia Wikipedia has accumulated 6.40 million articles and 250 million edits, contributed in a predominantly undirected and haphazard fashion by 5.77 million unvetted volunteers. Despite…

Digital Libraries · Computer Science 2007-05-23 Dennis M. Wilkinson , Bernardo A. Huberman

Writing Wikipedia with a neutral point of view is one of the five pillars of Wikipedia. Although the topic is core to Wikipedia, it is relatively understudied considering hundreds of research studies are published annually about the…

Computers and Society · Computer Science 2025-10-27 Isaac Johnson , Yu-Ming Liou , Jacob Rogers , Aaron Shaw , Leila Zia

Given Wikipedia's role as a trusted source of high-quality, reliable content, concerns are growing about the proliferation of low-quality machine-generated text (MGT) produced by large language models (LLMs) on its platform. Reliable…

Computation and Language · Computer Science 2025-07-08 Gerrit Quaremba , Elizabeth Black , Denny Vrandečić , Elena Simperl

Wikipedia's contents are based on reliable and published sources. To this date, relatively little is known about what sources Wikipedia relies on, in part because extracting citations and identifying cited sources is challenging. To close…

Digital Libraries · Computer Science 2020-11-24 Harshdeep Singh , Robert West , Giovanni Colavizza

Classification of bibliographic items into subjects and disciplines in large databases is essential for many quantitative science studies. The Web of Science classification of journals into ~250 subject categories, which has served as a…

Digital Libraries · Computer Science 2020-01-10 Staša Milojević

Content moderation in online platforms is crucial for ensuring activity therein adheres to existing policies, especially as these platforms grow. NLP research in this area has typically focused on automating some part of it given that it is…

Computation and Language · Computer Science 2024-08-13 Hsuvas Borkakoty , Luis Espinosa-Anke

In the past decade, the DBpedia community has put significant amount of effort on developing technical infrastructure and methods for efficient extraction of structured information from Wikipedia. These efforts have been primarily focused…

Computation and Language · Computer Science 2018-12-27 Milan Dojchinovski , Julio Hernandez , Markus Ackermann , Amit Kirschenbaum , Sebastian Hellmann
‹ Prev 1 8 9 10 Next ›