English
Related papers

Related papers: Automatic Wikipedia Link Generation Based On Inter…

200 papers

We are presenting a text analysis tool set that allows analysts in various fields to sieve through large collections of multilingual news items quickly and to find information that is of relevance to them. For a given document collection,…

Computation and Language · Computer Science 2007-05-23 Ralf Steinberger , Bruno Pouliquen , Camelia Ignat

There are over a billion websites on the Internet that can potentially serve as sources of information on various topics. One of the most popular examples of such an online source is Wikipedia. This public knowledge base is co-edited by…

Information Retrieval · Computer Science 2022-05-02 Włodzimierz Lewoniewski , Krzysztof Węcel , Witold Abramowicz

Wikipedia, the world largest encyclopedia contains a lot of knowledge that is expressed as formulae exclusively. Unfortunately, this knowledge is currently not fully accessible by intelligent information retrieval systems. This immense body…

Digital Libraries · Computer Science 2013-04-22 Moritz Schubotz

The large volume of scientific publications is likely to have hidden knowledge that can be used for suggesting new research topics. We propose an automatic method that is helpful for generating research hypotheses in the field of physics…

Information Retrieval · Computer Science 2017-12-27 Jung-Hun Kim , Aviv Segev

We present a tool that, from automatically recognised names, tries to infer inter-person relations in order to present associated people on maps. Based on an in-house Named Entity Recognition tool, applied on clusters of an average of…

Computation and Language · Computer Science 2007-05-23 Bruno Pouliquen , Ralf Steinberger , Camelia Ignat , Tamara Oellinger

We introduce MegaWika 2, a large, multilingual dataset of Wikipedia articles with their citations and scraped web sources; articles are represented in a rich data structure, and scraped source texts are stored inline with precise character…

Digital Libraries · Computer Science 2025-08-07 Samuel Barham , Chandler May , Benjamin Van Durme

The Internet-based encyclopaedia Wikipedia has grown to become one of the most visited web-sites on the Internet. However, critics have questioned the quality of entries, and an empirical study has shown Wikipedia to contain errors in a…

Digital Libraries · Computer Science 2011-01-04 Finn Aarup Nielsen

A simple dynamical model of collective edit activity of Wikipedia articles and their content evolution is introduced. Based on the recent empirical findings, each editor in the model is characterized by an ability to make content edit,…

Physics and Society · Physics 2023-04-25 Takashi Shimada , Fumiko Ogushi , Janos Torok , Janos Kertesz , Kimmo Kaski

Wikipedia is one of the most visited websites in the world and is also a frequent subject of scientific research. However, the analytical possibilities of Wikipedia information have not yet been analyzed considering at the same time both a…

Digital Libraries · Computer Science 2022-11-18 Wenceslao Arroyo-Machado , Daniel Torres-Salinas , Rodrigo Costas

This work improves monolingual sentence alignment for text simplification, specifically for text in standard and simple Wikipedia. We introduce a convolutional neural network structure to model similarity between two sentences. Due to the…

Computation and Language · Computer Science 2018-09-25 Yonghui Huang , Yunhui Li , Yi Luan

Geopolitics focuses on political power in relation to geographic space. Interactions among world countries have been widely studied at various scales, observing economic exchanges, world history or international politics among others. This…

Social and Information Networks · Computer Science 2017-07-21 Klaus M. Frahm , Samer El Zant , Katia Jaffrès-Runser , Dima L. Shepelyansky

Identifying which Wikipedia articles are related to science fiction, fantasy, or their hybrids is challenging because genre boundaries are porous and frequently overlap. Wikipedia nonetheless offers machine-readable structure beyond text,…

Information Retrieval · Computer Science 2026-03-02 Włodzimierz Lewoniewski , Milena Stróżyna , Izabela Czumałowska , Elżbieta Lewańska

Multilingualism is common offline, but we have a more limited understanding of the ways multilingualism is displayed online and the roles that multilinguals play in the spread of content between speakers of different languages. We take a…

Social and Information Networks · Computer Science 2016-06-14 Suin Kim , Sungjoon Park , Scott A. Hale , Sooyoung Kim , Jeongmin Byun , Alice Oh

In this paper, we present a comprehensive analysis and monitoring framework for the impact of Large Language Models (LLMs) on Wikipedia, examining the evolution of Wikipedia through existing data and using simulations to explore potential…

Computation and Language · Computer Science 2026-03-03 Siming Huang , Yuliang Xu , Mingmeng Geng , Yao Wan , Dongping Chen

We propose a dynamic map of knowledge generated from Wikipedia pages and the Web URLs contained therein. GalaxySearch provides answers to the questions we don't know how to ask, by constructing a semantic network of the most relevant pages…

Social and Information Networks · Computer Science 2012-04-17 Hauke Fuehres , Peter A. Gloor , Michael Henninger , Reto Kleeb , Keiichi Nemoto

We present Wikipedia-based Polyglot Dirichlet Allocation (WikiPDA), a crosslingual topic model that learns to represent Wikipedia articles written in any language as distributions over a common set of language-independent topics. It…

Computation and Language · Computer Science 2021-02-16 Tiziano Piccardi , Robert West

Verifiability is a core content policy of Wikipedia: claims that are likely to be challenged need to be backed by citations. There are millions of articles available online and thousands of new articles are released each month. For this…

As free online encyclopedias with massive volumes of content, Wikipedia and Wikidata are key to many Natural Language Processing (NLP) tasks, such as information retrieval, knowledge base building, machine translation, text classification,…

Using deep learning for different machine learning tasks such as image classification and word embedding has recently gained many attentions. Its appealing performance reported across specific Natural Language Processing (NLP) tasks in…

Computation and Language · Computer Science 2017-02-14 Ehsan Sherkat , Evangelos Milios

Wikipedia is the largest online encyclopedia: its open contribution policy allows everyone to edit and share their knowledge. A challenge of radical openness is that it facilitates introducing biased contents or perspectives in Wikipedia.…

Digital Libraries · Computer Science 2022-11-22 Puyu Yang , Giovanni Colavizza