中文
相关论文

相关论文: Wiki-Quantities and Wiki-Measurements: Datasets of…

200 篇论文

Wikipedia can be edited by anyone and thus contains various quality sentences. Therefore, Wikipedia includes some poor-quality edits, which are often marked up by other editors. While editors' reviews enhance the credibility of Wikipedia,…

计算与语言 · 计算机科学 2024-01-02 Kenichiro Ando , Satoshi Sekine , Mamoru Komachi

Wikipedia is the largest online encyclopedia, used by algorithms and web users as a central hub of reliable information on the web. The quality and reliability of Wikipedia content is maintained by a community of volunteer editors. Machine…

信息检索 · 计算机科学 2021-06-02 KayYen Wong , Miriam Redi , Diego Saez-Trumper

As free online encyclopedias with massive volumes of content, Wikipedia and Wikidata are key to many Natural Language Processing (NLP) tasks, such as information retrieval, knowledge base building, machine translation, text classification,…

Wikipedia's contents are based on reliable and published sources. To this date, relatively little is known about what sources Wikipedia relies on, in part because extracting citations and identifying cited sources is challenging. To close…

数字图书馆 · 计算机科学 2020-11-24 Harshdeep Singh , Robert West , Giovanni Colavizza

We introduce a new reading comprehension dataset, dubbed MultiWikiQA, which covers 306 languages and has 1,220,757 samples in total. We start with Wikipedia articles, which also provide the context for the dataset samples, and use an LLM to…

计算与语言 · 计算机科学 2026-03-05 Dan Saattrup Smart

Sequence-to-sequence models have recently gained the state of the art performance in summarization. However, not too many large-scale high-quality datasets are available and almost all the available ones are mainly news articles with…

计算与语言 · 计算机科学 2018-10-23 Mahnaz Koupaee , William Yang Wang

We present WikiReading, a large-scale natural language understanding task and publicly-available dataset with 18 million instances. The task is to predict textual values from the structured knowledge base Wikidata by reading the text of the…

We present a new dataset of Wikipedia articles each paired with a knowledge graph, to facilitate the research in conditional text generation, graph generation and graph representation learning. Existing graph-text paired datasets typically…

计算与语言 · 计算机科学 2021-07-21 Luyu Wang , Yujia Li , Ozlem Aslan , Oriol Vinyals

In this paper we present the Wikipedia Cultural Diversity dataset. For each existing Wikipedia language edition, the dataset contains a classification of the articles that represent its associated cultural context, i.e. all concepts and…

计算机与社会 · 计算机科学 2019-06-11 Marc Miquel-Ribé , David Laniado

We present a new concept - Wikiometrics - the derivation of metrics and indicators from Wikipedia. Wikipedia provides an accurate representation of the real world due to its size, structure, editing policy and popularity. We demonstrate an…

数字图书馆 · 计算机科学 2016-01-11 Gilad Katz , Lior Rokach

Wikipedia is an essential component of the open science ecosystem, yet it is poorly integrated with academic open science initiatives. Wikipedia Citations is a project that focuses on extracting and releasing comprehensive datasets of…

数字图书馆 · 计算机科学 2024-06-28 Natallia Kokash , Giovanni Colavizza

Datasets for data-to-text generation typically focus either on multi-domain, single-sentence generation or on single-domain, long-form generation. In this work, we cast generating Wikipedia sections as a data-to-text generation task and…

计算与语言 · 计算机科学 2021-06-03 Mingda Chen , Sam Wiseman , Kevin Gimpel

In this work, we open up the DAWT dataset - Densely Annotated Wikipedia Texts across multiple languages. The annotations include labeled text mentions mapping to entities (represented by their Freebase machine ids) as well as the type of…

信息检索 · 计算机科学 2017-03-06 Nemanja Spasojevic , Preeti Bhargava , Guoning Hu

With the growth of fake news and disinformation, the NLP community has been working to assist humans in fact-checking. However, most academic research has focused on model accuracy without paying attention to resource efficiency, which is…

计算机与社会 · 计算机科学 2021-09-03 Mykola Trokhymovych , Diego Saez-Trumper

We present a dataset that contains every instance of all tokens (~ words) ever written in undeleted, non-redirect English Wikipedia articles until October 2016, in total 13,545,349,787 instances. Each token is annotated with (i) the article…

计算与语言 · 计算机科学 2017-03-27 Fabian Flöck , Kenan Erdogan , Maribel Acosta

This study presents a comparative analysis of 55 Wikipedia language editions employing a citation index alongside a synthetic quality measure. Specifically, we identified the most significant Wikipedia articles within distinct topical…

信息检索 · 计算机科学 2025-05-23 Włodzimierz Lewoniewski , Krzysztof Węcel , Witold Abramowicz

Wikidata is one of the most important sources of structured data on the web, built by a worldwide community of volunteers. As a secondary source, its contents must be backed by credible references; this is particularly important as Wikidata…

人工智能 · 计算机科学 2021-09-21 Gabriel Amaral , Alessandro Piscopo , Lucie-Aimée Kaffee , Odinaldo Rodrigues , Elena Simperl

Wikidata is a collaborative knowledge graph which provides machine-readable structured data for Wikimedia projects including Wikipedia. Managed by a community of volunteers, it has grown to become the most edited Wikimedia project. However,…

社会与信息网络 · 计算机科学 2025-06-11 Marisa Ripoll , Neal Reeves , Anelia Kurteva , Elena Simperl , Albert Meroño Peñuela , Klaus Diepold

Wikidata has grown to a knowledge graph with an impressive size. To date, it contains more than 17 billion triples collecting information about people, places, films, stars, publications, proteins, and many more. On the other side, most of…

计算与语言 · 计算机科学 2024-01-17 Kunpeng Guo , Dennis Diefenbach , Antoine Gourru , Christophe Gravier

Wikidata is one of the most edited knowledge bases which contains structured data. It serves as the data source for many projects in the Wikimedia sphere and beyond. Since its inception in October 2012, it has been increasingly growing in…

数字图书馆 · 计算机科学 2019-11-19 Mariam Farda-Sarbas , Claudia Müller-Birn
‹ 上一页 1 2 3 10 下一页 ›