中文
相关论文

相关论文: Wiki Dumps to Training Corpora: South Slavic Case

200 篇论文

Automatic quality evaluation of Web information is a task with many fields of applications and of great relevance, especially in critical domains like the medical one. We move from the intuition that the quality of content of medical Web…

信息检索 · 计算机科学 2016-03-08 Vittoria Cozza , Marinella Petrocchi , Angelo Spognardi

This paper describes the corpus of sockpuppet cases we gathered from Wikipedia. A sockpuppet is an online user account created with a fake identity for the purpose of covering abusive behavior and/or subverting the editing regulation…

计算与语言 · 计算机科学 2013-10-28 Thamar Solorio , Ragib Hasan , Mainul Mizan

While Wikipedia has been utilized for fact-checking and claim verification to debunk misinformation and disinformation, it is essential to either improve article quality and rule out noisy articles. Self-contradiction is one of the…

计算与语言 · 计算机科学 2021-11-17 Cheng Hsu , Cheng-Te Li , Diego Saez-Trumper , Yi-Zhan Hsu

Online encyclopediae like Wikipedia contain large amounts of text that need frequent corrections and updates. The new information may contradict existing content in encyclopediae. In this paper, we focus on rewriting such dynamically…

计算与语言 · 计算机科学 2019-12-04 Darsh J Shah , Tal Schuster , Regina Barzilay

Wikipedia is among the largest examples of collective intelligence on the Web with over 61 million articles covering over 320 languages. Although edited and maintained by an active workforce of human volunteers, Wikipedia is highly reliant…

人机交互 · 计算机科学 2025-09-29 Neal Reeves , Elena Simperl

This paper describes a new system for semi-automatically building, extending and managing a terminological thesaurus---a multilingual terminology dictionary enriched with relationships between the terms themselves to form a thesaurus. The…

计算与语言 · 计算机科学 2019-04-09 Adam Rambousek , Ales Horak , Vit Suchomel , Vit Baisa

Wikidata is a collaborative knowledge graph which provides machine-readable structured data for Wikimedia projects including Wikipedia. Managed by a community of volunteers, it has grown to become the most edited Wikimedia project. However,…

社会与信息网络 · 计算机科学 2025-06-11 Marisa Ripoll , Neal Reeves , Anelia Kurteva , Elena Simperl , Albert Meroño Peñuela , Klaus Diepold

Wikipedia is a rich and invaluable source of information. Its central place on the Web makes it a particularly interesting object of study for scientists. Researchers from different domains used various complex datasets related to Wikipedia…

信息检索 · 计算机科学 2019-03-21 Nicolas Aspert , Volodymyr Miz , Benjamin Ricaud , Pierre Vandergheynst

In recent years, the field of document understanding has progressed a lot. A significant part of this progress has been possible thanks to the use of language models pretrained on large amounts of documents. However, pretraining corpora…

计算与语言 · 计算机科学 2023-06-07 Michał Turski , Tomasz Stanisławek , Karol Kaczmarek , Paweł Dyda , Filip Graliński

Wikipedia categories, a classification scheme built for organizing and describing Wikpedia articles, are being applied in computer science research. This paper adopts a systematic literature review approach, in order to identify different…

数字图书馆 · 计算机科学 2020-04-22 Jesús Tramullas , Piedad Garrido-Picazo , Ana I. Sánchez-Casabón

Systematic reviews, which entail the extraction of data from large numbers of scientific documents, are an ideal avenue for the application of machine learning. They are vital to many fields of science and philanthropy, but are very…

We compiled a new sentence splitting corpus that is composed of 203K pairs of aligned complex source and simplified target sentences. Contrary to previously proposed text simplification corpora, which contain only a small number of split…

计算与语言 · 计算机科学 2019-09-27 Christina Niklaus , Andre Freitas , Siegfried Handschuh

Corpora that contain tabular data such as WebTables are a vital resource for the academic community. Essentially, they are the backbone of any modern research in information management. They are used for various tasks of data extraction,…

计算与语言 · 计算机科学 2022-10-13 Platon Fedorov , Alexey Mironov , George Chernishev

Fake news detection is a challenging task aiming to reduce human time and effort to check the truthfulness of news. Automated approaches to combat fake news, however, are limited by the lack of labeled benchmark datasets, especially in…

计算与语言 · 计算机科学 2021-03-02 Inna Vogel , Jeong-Eun Choi , Meghana Meghana

Wikipedia serves as a globally accessible knowledge source with content in over 300 languages. Despite covering the same topics, the different versions of Wikipedia are written and updated independently. This leads to factual…

计算与语言 · 计算机科学 2026-05-19 Silvia Cappa , Lingxiao Kong , Pille-Riin Peet , Fanfu Wei , Yuchen Zhou , Jan-Christoph Kalo

This paper presents an approach to classify documents in any language into an English topical label space, without any text categorization training data. The approach, Cross-Lingual Dataless Document Classification (CLDDC) relies on mapping…

计算与语言 · 计算机科学 2016-11-15 Yangqiu Song , Stephen Mayhew , Dan Roth

Online encyclopedia such as Wikipedia has become one of the best sources of knowledge. Much effort has been devoted to expanding and enriching the structured data by automatic information extraction from unstructured text in Wikipedia.…

信息检索 · 计算机科学 2014-06-26 Kezun Zhang , Yanghua Xiao , Hanghang Tong , Haixun Wang , Wei Wang

We present a corpus that encompasses the complete history of conversations between contributors to Wikipedia, one of the largest online collaborative communities. By recording the intermediate states of conversations---including not only…

Large text corpora, such as Reddit posts, have become an increasingly prevalent site of qualitative inquiry. However, most large text corpora are intractable for qualitative researchers. Instead, teams rely on statistical subsampling to…

Text summarization is an approach for identifying important information present within text documents. This computational technique aims to generate shorter versions of the source text, by including only the relevant and salient information…

计算与语言 · 计算机科学 2021-06-30 Kalliath Abdul Rasheed Issam , Shivam Patel , Subalalitha C. N