English
Related papers

Related papers: Wiki Dumps to Training Corpora: South Slavic Case

200 papers

With the increasing demand for substantial amounts of high-quality data to train large language models (LLMs), efficiently filtering large web corpora has become a critical challenge. For this purpose, KenLM, a lightweight n-gram-based…

Computation and Language · Computer Science 2024-09-17 Yungi Kim , Hyunsoo Ha , Sukyung Lee , Jihoo Kim , Seonghoon Yang , Chanjun Park

The success of several architectures to learn semantic representations from unannotated text and the availability of these kind of texts in online multilingual resources such as Wikipedia has facilitated the massive and automatic creation…

Computation and Language · Computer Science 2020-03-31 Jesujoba O. Alabi , Kwabena Amponsah-Kaakyire , David I. Adelani , Cristina España-Bonet

The objective of the PANACEA ICT-2007.2.2 EU project is to build a platform that automates the stages involved in the acquisition, production, updating and maintenance of the large language resources required by, among others, MT systems.…

Computation and Language · Computer Science 2013-03-11 Núria Bel , Vassilis Papavasiliou , Prokopis Prokopidis , Antonio Toral , Victoria Arranz

Given Wikipedia's role as a trusted source of high-quality, reliable content, concerns are growing about the proliferation of low-quality machine-generated text (MGT) produced by large language models (LLMs) on its platform. Reliable…

Computation and Language · Computer Science 2025-07-08 Gerrit Quaremba , Elizabeth Black , Denny Vrandečić , Elena Simperl

We introduce a new reading comprehension dataset, dubbed MultiWikiQA, which covers 306 languages and has 1,220,757 samples in total. We start with Wikipedia articles, which also provide the context for the dataset samples, and use an LLM to…

Computation and Language · Computer Science 2026-03-05 Dan Saattrup Smart

It is now commonplace to observe that we are facing a deluge of online information. Researchers have of course long acknowledged the potential value of this information since digital traces make it possible to directly observe, describe and…

Computation and Language · Computer Science 2015-07-09 Thierry Poibeau , Pablo Ruiz

The explosion in the amount of news and journalistic content being generated across the globe, coupled with extended and instantaneous access to information through online media, makes it difficult and time-consuming to monitor news…

Computation and Language · Computer Science 2018-08-06 M. Tarik Altuncu , Sophia N. Yaliraki , Mauricio Barahona

In online communities, recent studies have strongly improved our knowledge about the different types or profiles of contributors, from casual to very involved ones, through focused people. However they do so by using very complex…

Human-Computer Interaction · Computer Science 2018-03-28 Shubham Krishna , Romain Billot , Nicolas Jullien

At a time when the quantity of - more or less freely - available data is increasing significantly, thanks to digital corpora, editions or libraries, the development of data mining tools or deep learning methods allows researchers to build a…

Computer Vision and Pattern Recognition · Computer Science 2019-04-29 Jean-Baptiste Camps , Gilles Guilhem Couffignal

Semantic textual similarity (STS) plays a crucial role in many natural language processing tasks. While extensively studied in high-resource languages, STS remains challenging for under-resourced languages such as Slovak. This paper…

Computation and Language · Computer Science 2026-02-05 Lukas Radosky , Miroslav Blstak , Matej Krajcovic , Ivan Polasek

Bilingual lexicons and phrase tables are critical resources for modern Machine Translation systems. Although recent results show that without any seed lexicon or parallel data, highly accurate bilingual lexicons can be learned using…

Computation and Language · Computer Science 2020-09-01 Ashiqur R. KhudaBukhsh , Shriphani Palakodety , Tom M. Mitchell

Constructing specialized content corpora from vast, unstructured web sources for domain-specific applications poses substantial data curation challenges. In this paper, we introduce a streamlined approach for generating high-quality,…

Computation and Language · Computer Science 2025-08-01 Franklin Zhang , Sonya Zhang , Alon Halevy

The proliferation of news media available online simultaneously presents a valuable resource and significant challenge to analysts aiming to profile and understand social and cultural trends in a geographic location of interest. While an…

Computation and Language · Computer Science 2021-08-18 A. Bock , A. Palladino , S. Smith-Heisters , I. Boardman , E. Pellegrini , E. J. Bienenstock , A. Valenti

Processing low-resource languages, such as Kiswahili, using machine learning is difficult due to lack of adequate training data. However, such low-resource languages are still important for human communication and are already in daily use…

Computation and Language · Computer Science 2025-01-17 Barack Wamkaya Wanjawa , Lawrence Muchemi , Evans Miriti

Automated terminology extraction refers to the task of extracting meaningful terms from domain-specific texts. This paper proposes a novel machine learning approach to terminology extraction, which combines features from traditional term…

Computation and Language · Computer Science 2025-02-25 Andraž Repar , Nada Lavrač , Senja Pollak

Recent research has taken advantage of Wikipedia's multilingualism as a resource for cross-language information retrieval and machine translation, as well as proposed techniques for enriching its cross-language structure. The availability…

Databases · Computer Science 2011-11-01 Thanh Nguyen , Viviane Moreira , Huong Nguyen , Hoa Nguyen , Juliana Freire

User generated text on social media often suffers from a lot of undesired characteristics including hatespeech, abusive language, insults etc. that are targeted to attack or abuse a specific group of people. Often such text is written…

Computation and Language · Computer Science 2019-10-03 Sravan Babu Bodapati , Spandana Gella , Kasturi Bhattacharjee , Yaser Al-Onaizan

Wikipedia's vision is a world in which everyone can share in the sum of all knowledge. In its first two decades, this vision has been very unevenly achieved. One of the largest hindrances is the sheer number of languages Wikipedia needs to…

Computers and Society · Computer Science 2020-04-13 Denny Vrandečić

While large language models (LLMs) can answer many questions correctly, they can also hallucinate and give wrong answers. Wikidata, with its over 12 billion facts, can be used to ground LLMs to improve their factuality. This paper presents…

Computation and Language · Computer Science 2023-11-07 Silei Xu , Shicheng Liu , Theo Culhane , Elizaveta Pertseva , Meng-Hsi Wu , Sina J. Semnani , Monica S. Lam

Knowledge Graphs are repositories of information that gather data from a multitude of domains and sources in the form of semantic triples, serving as a source of structured data for various crucial applications in the modern web landscape,…

Computation and Language · Computer Science 2022-10-27 Gabriel Amaral , Odinaldo Rodrigues , Elena Simperl
‹ Prev 1 8 9 10 Next ›