中文
相关论文

相关论文: WikiSQE: A Large-Scale Dataset for Sentence Qualit…

200 篇论文

Wikipedia serves as a globally accessible knowledge source with content in over 300 languages. Despite covering the same topics, the different versions of Wikipedia are written and updated independently. This leads to factual…

计算与语言 · 计算机科学 2026-05-19 Silvia Cappa , Lingxiao Kong , Pille-Riin Peet , Fanfu Wei , Yuchen Zhou , Jan-Christoph Kalo

Wikipedia is among the largest examples of collective intelligence on the Web with over 61 million articles covering over 320 languages. Although edited and maintained by an active workforce of human volunteers, Wikipedia is highly reliant…

人机交互 · 计算机科学 2025-09-29 Neal Reeves , Elena Simperl

We present a new dataset of Wikipedia articles each paired with a knowledge graph, to facilitate the research in conditional text generation, graph generation and graph representation learning. Existing graph-text paired datasets typically…

计算与语言 · 计算机科学 2021-07-21 Luyu Wang , Yujia Li , Ozlem Aslan , Oriol Vinyals

Wikidata is currently the largest open knowledge graph on the web, encompassing over 120 million entities. It integrates data from various domain-specific databases and imports a substantial amount of content from Wikipedia, while also…

计算与语言 · 计算机科学 2026-01-06 Shixiong Zhao , Hideaki Takeda

Query expansion (QE) is a well-known technique used to enhance the effectiveness of information retrieval. QE reformulates the initial query by adding similar terms that help in retrieving more relevant results. Several approaches have been…

信息检索 · 计算机科学 2019-06-21 Hiteshwar Kumar Azad , Akshay Deepak

Wikipedia is edited by volunteer editors around the world. Considering the large amount of existing content (e.g. over 5M articles in English Wikipedia), deciding what to edit next can be difficult, both for experienced users that usually…

信息检索 · 计算机科学 2020-09-25 Oleksii Moskalenko , Denis Parra , Diego Saez-Trumper

Text summarization is crucial for mitigating information overload across domains like journalism, medicine, and business. This research evaluates summarization performance across 17 large language models (OpenAI, Google, Anthropic,…

计算与语言 · 计算机科学 2025-04-08 Anantharaman Janakiraman , Behnaz Ghoraani

Today, comprehensive evaluation of large-scale machine learning models is possible thanks to the open datasets produced using crowdsourcing, such as SQuAD, MS COCO, ImageNet, SuperGLUE, etc. These datasets capture objective responses,…

人机交互 · 计算机科学 2021-11-29 Nikita Pavlichenko , Dmitry Ustalov

We present a new concept - Wikiometrics - the derivation of metrics and indicators from Wikipedia. Wikipedia provides an accurate representation of the real world due to its size, structure, editing policy and popularity. We demonstrate an…

数字图书馆 · 计算机科学 2016-01-11 Gilad Katz , Lior Rokach

Wikipedia is an online encyclopedia that anyone can edit. In this open model, some people edits with the intent of harming the integrity of Wikipedia. This is known as vandalism. We extend the framework presented in (Potthast, Stein, and…

信息检索 · 计算机科学 2012-10-23 Santiago M. Mola-Velasco

Our study identifies sentences in Wikipedia articles that are either identical or highly similar by applying techniques for near-duplicate detection of web pages. This is accomplished with a MapReduce implementation of minhash to identify…

信息检索 · 计算机科学 2014-06-05 Sarah Weissman , Samet Ayhan , Joshua Bradley , Jimmy Lin

As large language models (LLMs) grow larger and more sophisticated, assessing their "reasoning" capabilities in natural language grows more challenging. Recent question answering (QA) benchmarks that attempt to assess reasoning are often…

计算与语言 · 计算机科学 2022-12-01 Matthew Ho , Aditya Sharma , Justin Chang , Michael Saxon , Sharon Levy , Yujie Lu , William Yang Wang

This paper presents TextComplexityDE, a dataset consisting of 1000 sentences in German language taken from 23 Wikipedia articles in 3 different article-genres to be used for developing text-complexity predictor models and automatic text…

计算与语言 · 计算机科学 2019-04-17 Babak Naderi , Salar Mohtaj , Kaspar Ensikat , Sebastian Möller

News article revision histories provide clues to narrative and factual evolution in news articles. To facilitate analysis of this evolution, we present the first publicly available dataset of news revision histories, NewsEdits. Our dataset…

计算与语言 · 计算机科学 2022-06-16 Alexander Spangher , Xiang Ren , Jonathan May , Nanyun Peng

Wikidata is a collaborative knowledge graph which provides machine-readable structured data for Wikimedia projects including Wikipedia. Managed by a community of volunteers, it has grown to become the most edited Wikimedia project. However,…

社会与信息网络 · 计算机科学 2025-06-11 Marisa Ripoll , Neal Reeves , Anelia Kurteva , Elena Simperl , Albert Meroño Peñuela , Klaus Diepold

Visual Question Answering (VQA) benchmarks have largely emphasized perception-based tasks that can be solved from visual content alone. In contrast, many real-world scenarios require external knowledge that is not directly observable in the…

计算机视觉与模式识别 · 计算机科学 2026-05-21 Basel Shbita , Pengyuan Li , Anna Lisa Gentile

A simple dynamical model of collective edit activity of Wikipedia articles and their content evolution is introduced. Based on the recent empirical findings, each editor in the model is characterized by an ability to make content edit,…

物理与社会 · 物理学 2023-04-25 Takashi Shimada , Fumiko Ogushi , Janos Torok , Janos Kertesz , Kimmo Kaski

To foster the development of new models for collaborative AI-assisted report generation, we introduce MegaWika, consisting of 13 million Wikipedia articles in 50 diverse languages, along with their 71 million referenced source materials. We…

Wikidata is a multi-language knowledge base that is being edited and maintained by editors from different language communities. Due to the structured nature of its content, the contributions are in various forms, including manual edit,…

人机交互 · 计算机科学 2023-11-07 Jeffrey Jun-jie Ma , Charles Chuankai Zhang

Hoaxes are a recognised form of disinformation created deliberately, with potential serious implications in the credibility of reference knowledge resources such as Wikipedia. What makes detecting Wikipedia hoaxes hard is that they often…

计算与语言 · 计算机科学 2024-09-02 Hsuvas Borkakoty , Luis Espinosa-Anke