English
Related papers

Related papers: WikiDoMiner: Wikipedia Domain-specific Miner

200 papers

English Wikipedia has long been an important data source for much research and natural language machine learning modeling. The growth of non-English language editions of Wikipedia, greater computational resources, and calls for equity in…

Computers and Society · Computer Science 2022-04-07 Isaac Johnson , Emily Lescak

One prerequisite for supervised machine learning is high quality labelled data. Acquiring such data is, particularly if expert knowledge is required, costly or even impossible if the task needs to be performed by a single expert. In this…

Software Engineering · Computer Science 2023-10-02 Michael Unterkalmsteiner , Andrew Yates

A fundamental challenge in the current NLP context, dominated by language models, comes from the inflexibility of current architectures to 'learn' new information. While model-centric solutions like continual learning or parameter-efficient…

Computation and Language · Computer Science 2023-08-21 Hsuvas Borkakoty , Luis Espinosa-Anke

There are over a billion websites on the Internet that can potentially serve as sources of information on various topics. One of the most popular examples of such an online source is Wikipedia. This public knowledge base is co-edited by…

Information Retrieval · Computer Science 2022-05-02 Włodzimierz Lewoniewski , Krzysztof Węcel , Witold Abramowicz

Wikipedia articles contain multiple links connecting a subject to other pages of the encyclopedia. In Wikipedia parlance, these links are called internal links or wikilinks. We present a complete dataset of the network of internal Wikipedia…

Social and Information Networks · Computer Science 2019-04-05 Cristian Consonni , David Laniado , Alberto Montresor

References are an essential part of Wikipedia. Each statement in Wikipedia should be referenced. In this paper, we explore the creation and collection of references for new Wikipedia articles from an editors' perspective. We map out the…

Computation and Language · Computer Science 2021-03-05 Lucie-Aimée Kaffee , Hady Elsahar

Since its emergence in the 1990s the World Wide Web (WWW) has rapidly evolved into a huge mine of global information and it is growing in size everyday. The presence of huge amount of resources on the Web thus poses a serious problem of…

Information Retrieval · Computer Science 2011-02-04 Debajyoti Mukhopadhyay , Aritra Banik , Sreemoyee Mukherjee , Jhilik Bhattacharya , Young-Chon Kim

Wikidata is a frequently updated, community-driven, and multilingual knowledge graph. Hence, Wikidata is an attractive basis for Entity Linking, which is evident by the recent increase in published papers. This survey focuses on four…

Computation and Language · Computer Science 2021-12-06 Cedric Möller , Jens Lehmann , Ricardo Usbeck

We present HowSumm, a novel large-scale dataset for the task of query-focused multi-document summarization (qMDS), which targets the use-case of generating actionable instructions from a set of sources. This use-case is different from the…

Computation and Language · Computer Science 2021-10-12 Odellia Boni , Guy Feigenblat , Guy Lev , Michal Shmueli-Scheuer , Benjamin Sznajder , David Konopnicki

The rapidly increasing number of scientific documents available publicly on the Internet creates the challenge of efficiently organizing and indexing these documents. Due to the time consuming and tedious nature of manual classification and…

Digital Libraries · Computer Science 2016-03-22 Lisa Posch

In online communities, recent studies have strongly improved our knowledge about the different types or profiles of contributors, from casual to very involved ones, through focused people. However they do so by using very complex…

Human-Computer Interaction · Computer Science 2018-03-28 Shubham Krishna , Romain Billot , Nicolas Jullien

We present a configurable pipeline for generating multilingual sets of entities with specified characteristics, such as domain, geographical location and popularity, using data from Wikipedia and Wikidata. These datasets are intended for…

Computation and Language · Computer Science 2026-04-02 Pavel Braslavski , Dmitrii Iarosh , Nikita Sushko , Andrey Sakhovskiy , Vasily Konovalov , Elena Tutubalina , Alexander Panchenko

The Wikimedia Foundation has recently observed that newly joining editors on Wikipedia are increasingly failing to integrate into the Wikipedia editors' community, i.e. the community is becoming increasingly harder to penetrate. To sustain…

Computers and Society · Computer Science 2014-05-30 Kalpit V Desai , Roopesh Ranjan

Wikidata is one of the most edited knowledge bases which contains structured data. It serves as the data source for many projects in the Wikimedia sphere and beyond. Since its inception in October 2012, it has been increasingly growing in…

Digital Libraries · Computer Science 2019-11-19 Mariam Farda-Sarbas , Claudia Müller-Birn

Users of spoken dialogue systems (SDS) expect high quality interactions across a wide range of diverse topics. However, the implementation of SDS capable of responding to every conceivable user utterance in an informative way is a…

Computation and Language · Computer Science 2021-01-28 A. Augustin , A. Papangelis , M. Kotti , P. Vougiouklis , J. Hare , N. Braunschweiler

Wikipedia is written in the wikitext markup language. When serving content, the MediaWiki software that powers Wikipedia parses wikitext to HTML, thereby inserting additional content by expanding macros (templates and mod-ules). Hence,…

Computers and Society · Computer Science 2020-04-22 Blagoj Mitrevski , Tiziano Piccardi , Robert West

Knowledge graphs (KGs) have become ubiquitous publicly available knowledge sources, and are nowadays covering an ever increasing array of domains. However, not all knowledge represented is useful or pertaining when considering a new…

Artificial Intelligence · Computer Science 2024-10-18 Pierre Monnin , Cherif-Hassan Nousradine , Lucas Jarnac , Laurel Zuckerman , Miguel Couceiro

Parallel sentences are a relatively scarce but extremely useful resource for many applications including cross-lingual retrieval and statistical machine translation. This research explores our new methodologies for mining such data from…

Computation and Language · Computer Science 2015-11-20 Krzysztof Wołk , Emilia Rejmund , Krzysztof Marasek

Question Answering (QA) is a growing area of research, often used to facilitate the extraction of information from within documents. State-of-the-art QA models are usually pre-trained on domain-general corpora like Wikipedia and thus tend…

Computation and Language · Computer Science 2022-12-01 Matthew Maufe , James Ravenscroft , Rob Procter , Maria Liakata

The Wikipedia category graph serves as the taxonomic backbone for large-scale knowledge graphs like YAGO or Probase, and has been used extensively for tasks like entity disambiguation or semantic similarity estimation. Wikipedia's…

Information Retrieval · Computer Science 2019-07-01 Nicolas Heist , Heiko Paulheim
‹ Prev 1 8 9 10 Next ›