中文
相关论文

相关论文: MegaWika 2: A More Comprehensive Multilingual Coll…

200 篇论文

To foster the development of new models for collaborative AI-assisted report generation, we introduce MegaWika, consisting of 13 million Wikipedia articles in 50 diverse languages, along with their 71 million referenced source materials. We…

Wikipedia is an essential component of the open science ecosystem, yet it is poorly integrated with academic open science initiatives. Wikipedia Citations is a project that focuses on extracting and releasing comprehensive datasets of…

数字图书馆 · 计算机科学 2024-06-28 Natallia Kokash , Giovanni Colavizza

Wikipedia is the largest online encyclopedia, used by algorithms and web users as a central hub of reliable information on the web. The quality and reliability of Wikipedia content is maintained by a community of volunteer editors. Machine…

信息检索 · 计算机科学 2021-06-02 KayYen Wong , Miriam Redi , Diego Saez-Trumper

Wikipedia is a critical source of information for millions of users across the Web. It serves as a key resource for large language models, search engines, question-answering systems, and other Web-based applications. In Wikipedia, content…

Wikipedia articles representing an entity or a topic in different language editions evolve independently within the scope of the language-specific user communities. This can lead to different points of views reflected in the articles, as…

计算与语言 · 计算机科学 2017-02-03 Simon Gottschalk , Elena Demidova

With over 60M articles, Wikipedia has become the largest platform for open and freely accessible knowledge. While it has more than 15B monthly visits, its content is believed to be inaccessible to many readers due to the lack of readability…

计算与语言 · 计算机科学 2024-06-05 Mykola Trokhymovych , Indira Sen , Martin Gerlach

Wikidata is one of the most edited knowledge bases which contains structured data. It serves as the data source for many projects in the Wikimedia sphere and beyond. Since its inception in October 2012, it has been increasingly growing in…

数字图书馆 · 计算机科学 2019-11-19 Mariam Farda-Sarbas , Claudia Müller-Birn

Wikipedia's contents are based on reliable and published sources. To this date, relatively little is known about what sources Wikipedia relies on, in part because extracting citations and identifying cited sources is challenging. To close…

数字图书馆 · 计算机科学 2020-11-24 Harshdeep Singh , Robert West , Giovanni Colavizza

The different Wikipedia language editions vary dramatically in how comprehensive they are. As a result, most language editions contain only a small fraction of the sum of information that exists across all Wikipedias. In this paper, we…

社会与信息网络 · 计算机科学 2016-04-13 Ellery Wulczyn , Robert West , Leila Zia , Jure Leskovec

INTRODUCTION: Wikipedia is a major source of information, particularly for medical and health content, citing over 4 million scholarly publications. However, the representation of research-based knowledge across different languages on…

数字图书馆 · 计算机科学 2025-01-17 Michael Taylor , Roisi Proven , Carlos Areia

Fast-developing fields such as Artificial Intelligence (AI) often outpace the efforts of encyclopedic sources such as Wikipedia, which either do not completely cover recently-introduced topics or lack such content entirely. As a result,…

An important editing policy in Wikipedia is to provide citations for added statements in Wikipedia pages, where statements can be arbitrary pieces of text, ranging from a sentence to a paragraph. In many cases citations are either outdated…

信息检索 · 计算机科学 2017-04-26 Besnik Fetahu , Katja Markert , Wolfgang Nejdl , Avishek Anand

We introduce a new reading comprehension dataset, dubbed MultiWikiQA, which covers 306 languages and has 1,220,757 samples in total. We start with Wikipedia articles, which also provide the context for the dataset samples, and use an LLM to…

计算与语言 · 计算机科学 2026-03-05 Dan Saattrup Smart

We introduce WikiLingua, a large-scale, multilingual dataset for the evaluation of crosslingual abstractive summarization systems. We extract article and summary pairs in 18 languages from WikiHow, a high quality, collaborative resource of…

计算与语言 · 计算机科学 2020-10-08 Faisal Ladhak , Esin Durmus , Claire Cardie , Kathleen McKeown

Webpages have been a rich resource for language and vision-language tasks. Yet only pieces of webpages are kept: image-caption pairs, long text articles, or raw HTML, never all in one place. Webpage tasks have resultingly received little…

计算与语言 · 计算机科学 2023-05-10 Andrea Burns , Krishna Srinivasan , Joshua Ainslie , Geoff Brown , Bryan A. Plummer , Kate Saenko , Jianmo Ni , Mandy Guo

Online encyclopediae like Wikipedia contain large amounts of text that need frequent corrections and updates. The new information may contradict existing content in encyclopediae. In this paper, we focus on rewriting such dynamically…

计算与语言 · 计算机科学 2019-12-04 Darsh J Shah , Tal Schuster , Regina Barzilay

With more than 11 times as many pageviews as the next largest edition, English Wikipedia dominates global knowledge access relative to other language editions. Readers are prone to assuming English Wikipedia as a superset of all language…

人机交互 · 计算机科学 2026-01-21 Zining Wang , Yuxuan Zhang , Dongwook Yoon , Nicholas Vincent , Farhan Samir , Vered Shwartz

In this paper, we present a dataset of inter-language knowledge propagation in Wikipedia. Covering the entire 309 language editions and 33M articles, the dataset aims to track the full propagation history of Wikipedia concepts, and allow…

计算机与社会 · 计算机科学 2021-04-01 Roldolfo Valentim , Giovanni Comarela , Souneil Park , Diego Saez-Trumper

Wikipedia serves as a globally accessible knowledge source with content in over 300 languages. Despite covering the same topics, the different versions of Wikipedia are written and updated independently. This leads to factual…

计算与语言 · 计算机科学 2026-05-19 Silvia Cappa , Lingxiao Kong , Pille-Riin Peet , Fanfu Wei , Yuchen Zhou , Jan-Christoph Kalo

Webpages have been a rich, scalable resource for vision-language and language only tasks. Yet only pieces of webpages are kept in existing datasets: image-caption pairs, long text articles, or raw HTML, never all in one place. Webpage tasks…

计算与语言 · 计算机科学 2023-10-23 Andrea Burns , Krishna Srinivasan , Joshua Ainslie , Geoff Brown , Bryan A. Plummer , Kate Saenko , Jianmo Ni , Mandy Guo
‹ 上一页 1 2 3 10 下一页 ›