English
Related papers

Related papers: Wiki Dumps to Training Corpora: South Slavic Case

200 papers

Well curated, large-scale corpora of social media posts containing broad public opinion offer an alternative data source to complement traditional surveys. While surveys are effective at collecting representative samples and are capable of…

Computation and Language · Computer Science 2025-02-14 Michael V. Arnold , Peter Sheridan Dodds , Christopher M. Danforth

Nowadays, thanks to Web 2.0 technologies, people have the possibility to generate and spread contents on different social media in a very easy way. In this context, the evaluation of the quality of the information that is available online…

Computation and Language · Computer Science 2018-12-10 Elias Bassani , Marco Viviani

Wikipedia is a critical source of information for millions of users across the Web. It serves as a key resource for large language models, search engines, question-answering systems, and other Web-based applications. In Wikipedia, content…

The availability of parallel sentence simplification (SS) is scarce for neural SS modelings. We propose an unsupervised method to build SS corpora from large-scale bilingual translation corpora, alleviating the need for SS supervised…

Computation and Language · Computer Science 2021-09-02 Xinyu Lu , Jipeng Qiang , Yun Li , Yunhao Yuan , Yi Zhu

Wikimedia content is used extensively by the AI community and within the language modeling community in particular. In this paper, we provide a review of the different ways in which Wikimedia data is curated to use in NLP tasks across…

Computers and Society · Computer Science 2024-10-14 Isaac Johnson , Lucie-Aimée Kaffee , Miriam Redi

Wikidata is a free and open knowledge base from the Wikimedia Foundation, that not only acts as a central storage of structured data for other projects of the organization, but also for a growing array of information systems, including…

We present Wikipedia-based Polyglot Dirichlet Allocation (WikiPDA), a crosslingual topic model that learns to represent Wikipedia articles written in any language as distributions over a common set of language-independent topics. It…

Computation and Language · Computer Science 2021-02-16 Tiziano Piccardi , Robert West

Naturally-occurring instances of linguistic phenomena are important both for training and for evaluating automatic processes on text. When available in large quantities, they also prove interesting material for linguistic studies. In this…

Computation and Language · Computer Science 2022-02-28 Aurélien Max , Guillaume Wisniewski

We test the hypothesis that the extent to which one obtains information on a given topic through Wikipedia depends on the language in which it is consulted. Controlling the size factor, we investigate this hypothesis for a number of 25…

Computation and Language · Computer Science 2021-06-01 Alexander Mehler , Wahed Hemati , Pascal Welke , Maxim Konca , Tolga Uslu

With over 60M articles, Wikipedia has become the largest platform for open and freely accessible knowledge. While it has more than 15B monthly visits, its content is believed to be inaccessible to many readers due to the lack of readability…

Computation and Language · Computer Science 2024-06-05 Mykola Trokhymovych , Indira Sen , Martin Gerlach

In this paper we present a profile-based approach to information filtering by an analysis of the content of text documents. The Wikipedia index database is created and used to automatically generate the user profile from the user document…

Information Retrieval · Computer Science 2008-05-08 A. V. Smirnov , A. A. Krizhanovsky

The algorithm of the creation texts parallel corpora was presented. The algorithm is based on the use of "key words" in text documents, and on the means of their automated translation. Key words were singled out by means of using Russian…

Computation and Language · Computer Science 2008-07-03 D. V. Lande , V. V. Zhygalo

We introduce GeBioToolkit, a tool for extracting multilingual parallel corpora at sentence level, with document and gender information from Wikipedia biographies. Despite thegender inequalitiespresent in Wikipedia, the toolkit has been…

Computation and Language · Computer Science 2019-12-11 Marta R. Costa-jussà , Pau Li Lin , Cristina España-Bonet

This paper investigates the impact of corpus creation decisions on large multi-lingual geographic web corpora. Beginning with a 427 billion word corpus derived from the Common Crawl, three methods are used to improve the quality of…

Computation and Language · Computer Science 2024-03-14 Jonathan Dunn

Debate portals and similar web platforms constitute one of the main text sources in computational argumentation research and its applications. While the corpora built upon these sources are rich of argumentatively relevant content and…

Computation and Language · Computer Science 2020-11-04 Jonas Dorsch , Henning Wachsmuth

Over the last few years, verifying the credibility of information sources has become a fundamental need to combat disinformation. Here, we present a language-agnostic model designed to assess the reliability of web domains as sources in…

Social and Information Networks · Computer Science 2025-11-21 Jacopo D'Ignazi , Andreas Kaltenbrunner , Yelena Mejova , Michele Tizzani , Kyriaki Kalimeri , Mariano Beiró , Pablo Aragón

Malicious sockpuppet detection on Wikipedia is critical to preserving access to reliable information on the internet and preventing the spread of disinformation. Prior machine learning approaches rely on stylistic and meta-data features,…

Machine Learning · Computer Science 2025-10-29 Luc Raszewski , Christine De Kock

Large language models have led to remarkable progress on many NLP tasks, and researchers are turning to ever-larger text corpora to train them. Some of the largest corpora available are made by scraping significant portions of the internet,…

Computation and Language · Computer Science 2021-10-01 Jesse Dodge , Maarten Sap , Ana Marasović , William Agnew , Gabriel Ilharco , Dirk Groeneveld , Margaret Mitchell , Matt Gardner

Hierarchical domain-specific classification schemas (or subject heading vocabularies) are often used to identify, classify, and disambiguate concepts that occur in scholarly articles. In this work, we develop, apply, and evaluate a…

Social and Information Networks · Computer Science 2021-09-13 Kanyao Han , Pingjing Yang , Shubhanshu Mishra , Jana Diesner

Wikipedia is the world's largest online encyclopedia, but maintaining article quality through collaboration is challenging. Wikipedia designed a quality scale, but with such a manual assessment process, many articles remain unassessed. We…

Computation and Language · Computer Science 2023-10-04 Pedro Miguel Moás , Carla Teixeira Lopes