English
Related papers

Related papers: Wiki Dumps to Training Corpora: South Slavic Case

200 papers

Fast-developing fields such as Artificial Intelligence (AI) often outpace the efforts of encyclopedic sources such as Wikipedia, which either do not completely cover recently-introduced topics or lack such content entirely. As a result,…

Computation and Language · Computer Science 2022-06-23 Irene Li , Alexander Fabbri , Rina Kawamura , Yixin Liu , Xiangru Tang , Jaesung Tae , Chang Shen , Sally Ma , Tomoe Mizutani , Dragomir Radev

English Wikipedia has long been an important data source for much research and natural language machine learning modeling. The growth of non-English language editions of Wikipedia, greater computational resources, and calls for equity in…

Computers and Society · Computer Science 2022-04-07 Isaac Johnson , Emily Lescak

This work compares concept models for cross-language retrieval: First, we adapt probabilistic Latent Semantic Analysis (pLSA) for multilingual documents. Experiments with different weighting schemes show that a weighting method favoring…

Information Retrieval · Computer Science 2014-01-13 Benjamin Roth

The task of discovering topics in text corpora has been dominated by Latent Dirichlet Allocation and other Topic Models for over a decade. In order to apply these approaches to massive text corpora, the vocabulary needs to be reduced…

Computation and Language · Computer Science 2019-08-08 Gibran Fuentes-Pineda , Ivan Vladimir Meza-Ruiz

The creation of systematic literature reviews (SLR) is critical for analyzing the landscape of a research field and guiding future research directions. However, retrieving and filtering the literature corpus for an SLR is highly…

Machine Learning · Computer Science 2026-02-18 Lucas Joos , Daniel A. Keim , Maximilian T. Fischer

Highly specific datasets of scientific literature are important for both research and education. However, it is difficult to build such datasets at scale. A common approach is to build these datasets reductively by applying topic modeling…

Information Retrieval · Computer Science 2023-09-20 Nicholas Solovyev , Ryan Barron , Manish Bhattarai , Maksim E. Eren , Kim O. Rasmussen , Boian S. Alexandrov

Most people do not interact with Semantic Web data directly. Unless they have the expertise to understand the underlying technology, they need textual or visual interfaces to help them make sense of it. We explore the problem of generating…

Computation and Language · Computer Science 2024-07-04 Pavlos Vougiouklis , Hady Elsahar , Lucie-Aimée Kaffee , Christoph Gravier , Frederique Laforest , Jonathon Hare , Elena Simperl

A Wikipedia book (known as Wikibook) is a collection of Wikipedia articles on a particular theme that is organized as a book. We propose Wikibook-Bot, a machine-learning based technique for automatically generating high quality Wikibooks…

Digital Libraries · Computer Science 2018-12-31 Shahar Admati , Lior Rokach , Bracha Shapira

In recent years, research on visual document understanding (VDU) has grown significantly, with a particular emphasis on the development of self-supervised learning methods. However, one of the significant challenges faced in this field is…

Computer Vision and Pattern Recognition · Computer Science 2023-05-03 Donghyun Kim , Teakgyu Hong , Moonbin Yim , Yoonsik Kim , Geewook Kim

The task of Split and Rephrase, which splits a complex sentence into multiple simple sentences with the same meaning, improves readability and enhances the performance of downstream tasks in natural language processing (NLP). However, while…

Computation and Language · Computer Science 2024-04-16 Hayato Tsukagoshi , Tsutomu Hirao , Makoto Morishita , Katsuki Chousa , Ryohei Sasano , Koichi Takeda

We propose an edit-centric approach to assess Wikipedia article quality as a complementary alternative to current full document-based techniques. Our model consists of a main classifier equipped with an auxiliary generative module which,…

Computation and Language · Computer Science 2019-09-20 Edison Marrese-Taylor , Pablo Loyola , Yutaka Matsuo

We introduce WikiLingua, a large-scale, multilingual dataset for the evaluation of crosslingual abstractive summarization systems. We extract article and summary pairs in 18 languages from WikiHow, a high quality, collaborative resource of…

Computation and Language · Computer Science 2020-10-08 Faisal Ladhak , Esin Durmus , Claire Cardie , Kathleen McKeown

Text editing, i.e., the process of modifying or manipulating text, is a crucial step in human writing process. In this paper, we study the problem of controlled text editing by natural language instruction. According to a given instruction…

Computation and Language · Computer Science 2023-10-10 Xiang Chen , Zheng Li , Xiaojun Wan

One of the most impressive human endeavors of the past two decades is the collection and categorization of human knowledge in the free and accessible format that is Wikipedia. In this work we ask what makes a term worthy of entering this…

Computation and Language · Computer Science 2020-09-18 Yonatan Bilu , Shai Gretz , Edo Cohen , Noam Slonim

Machine learning applications to symbolic mathematics are becoming increasingly popular, yet there lacks a centralized source of real-world symbolic expressions to be used as training data. In contrast, the field of natural language…

Machine Learning · Computer Science 2022-07-06 Joanne T. Kim , Mikel Landajuela , Brenden K. Petersen

The different Wikipedia language editions vary dramatically in how comprehensive they are. As a result, most language editions contain only a small fraction of the sum of information that exists across all Wikipedias. In this paper, we…

Social and Information Networks · Computer Science 2016-04-13 Ellery Wulczyn , Robert West , Leila Zia , Jure Leskovec

Wikipedia, a paradigmatic example of online knowledge space is organized in a collaborative, bottom-up way with voluntary contributions, yet it maintains a level of reliability comparable to that of traditional encyclopedias. The lack of…

Physics and Society · Physics 2021-05-24 Fumiko Ogushi , János Kertész , Kimmo Kaski , Takashi Shimada

Wikidata has been increasingly adopted by many communities for a wide variety of applications, which demand high-quality knowledge to deliver successful results. In this paper, we develop a framework to detect and analyze low-quality…

Artificial Intelligence · Computer Science 2021-11-22 Kartik Shenoy , Filip Ilievski , Daniel Garijo , Daniel Schwabe , Pedro Szekely

Cross-lingual summarization (CLS) is the task to produce a summary in one particular language for a source document in a different language. We introduce WikiMulti - a new dataset for cross-lingual summarization based on Wikipedia articles…

Computation and Language · Computer Science 2022-04-26 Pavel Tikhonov , Valentin Malykh

Wikipedia can easily be justified as a behemoth, considering the sheer volume of content that is added or removed every minute to its several projects. This creates an immense scope, in the field of natural language processing towards…

Social and Information Networks · Computer Science 2019-06-12 Soumya Sarkar , Bhanu Prakash Reddy , Sandipan Sikdar , Animesh Mukherjee
‹ Prev 1 3 4 5 6 7 10 Next ›