中文
相关论文

相关论文: WikiSQE: A Large-Scale Dataset for Sentence Qualit…

200 篇论文

FActScore has gained popularity as a metric to estimate the factuality of long-form texts generated by Large Language Models (LLMs) in English. However, there has not been any work in studying the behavior of FActScore in other languages.…

计算与语言 · 计算机科学 2024-07-01 Kim Trong Vu , Michael Krumdick , Varshini Reddy , Franck Dernoncourt , Viet Dac Lai

With this work, we present a publicly available dataset of the history of all the references (more than 55 million) ever used in the English Wikipedia until June 2019. We have applied a new method for identifying and monitoring references…

计算机与社会 · 计算机科学 2020-10-08 Olga Zagovora , Roberto Ulloa , Katrin Weller , Fabian Flöck

To identify and classify toxic online commentary, the modern tools of data science transform raw text into key features from which either thresholding or learning algorithms can make predictions for monitoring offensive conversations. We…

机器学习 · 计算机科学 2018-10-05 David Noever

We study text reuse related to Wikipedia at scale by compiling the first corpus of text reuse cases within Wikipedia as well as without (i.e., reuse of Wikipedia text in a sample of the Common Crawl). To discover reuse beyond verbatim copy…

Among the manifold takes on world literature, it is our goal to contribute to the discussion from a digital point of view by analyzing the representation of world literature in Wikipedia with its millions of articles in hundreds of…

信息检索 · 计算机科学 2017-01-05 Christoph Hube , Frank Fischer , Robert Jäschke , Gerhard Lauer , Mads Rosendahl Thomsen

We present the Stanford Question Answering Dataset (SQuAD), a new reading comprehension dataset consisting of 100,000+ questions posed by crowdworkers on a set of Wikipedia articles, where the answer to each question is a segment of text…

计算与语言 · 计算机科学 2016-10-12 Pranav Rajpurkar , Jian Zhang , Konstantin Lopyrev , Percy Liang

We investigate the generation of one-sentence Wikipedia biographies from facts derived from Wikidata slot-value pairs. We train a recurrent neural network sequence-to-sequence model with attention to select facts and generate textual…

计算与语言 · 计算机科学 2017-02-22 Andrew Chisholm , Will Radford , Ben Hachey

Wikidata has grown to a knowledge graph with an impressive size. To date, it contains more than 17 billion triples collecting information about people, places, films, stars, publications, proteins, and many more. On the other side, most of…

计算与语言 · 计算机科学 2024-01-17 Kunpeng Guo , Dennis Diefenbach , Antoine Gourru , Christophe Gravier

Wikipedia has high-quality articles on a variety of topics and has been used in diverse research areas. In this study, a method is presented for using Wikipedia's editor information to build recommender systems in various domains that…

信息检索 · 计算机科学 2023-06-16 Katsuhiko Hayashi

Tabular data, as a crucial form of data representation, exists in diverse formats on the Web. When confronted with complex and irregular tables, manual modification becomes a laborious task. This paper investigates the performance of Large…

人工智能 · 计算机科学 2024-03-06 Zheng Li , Xiang Chen , Xiaojun Wan

Wikimedia content is used extensively by the AI community and within the language modeling community in particular. In this paper, we provide a review of the different ways in which Wikimedia data is curated to use in NLP tasks across…

计算机与社会 · 计算机科学 2024-10-14 Isaac Johnson , Lucie-Aimée Kaffee , Miriam Redi

Predicting which words are considered hard to understand for a given target population is a vital step in many NLP applications such as text simplification. This task is commonly referred to as Complex Word Identification (CWI). With a few…

计算与语言 · 计算机科学 2020-06-12 Matthew Shardlow , Michael Cooper , Marcos Zampieri

The performance and usability of Large-Language Models (LLMs) are driving their use in explanation generation tasks. However, despite their widespread adoption, LLM explanations have been found to be unreliable, making it difficult for…

While Wikipedia exists in 287 languages, its content is unevenly distributed among them. In this work, we investigate the generation of open domain Wikipedia summaries in underserved languages using structured data from Wikidata. To this…

Documents are fundamental to preserving and disseminating information, often incorporating complex layouts, tables, and charts that pose significant challenges for automatic document understanding (DU). While vision-language large models…

计算与语言 · 计算机科学 2025-06-19 Negar Foroutan , Angelika Romanou , Matin Ansaripour , Julian Martin Eisenschlos , Karl Aberer , Rémi Lebret

The Internet has provided us with great opportunities for large scale collaborative public good projects. Wikipedia is a predominant example of such projects where conflicts emerge and get resolved through bottom-up mechanisms leading to…

物理与社会 · 物理学 2017-06-30 Csilla Rudas , Olivér Surányi , Taha Yasseri , János Török

In this paper, we present a comprehensive analysis and monitoring framework for the impact of Large Language Models (LLMs) on Wikipedia, examining the evolution of Wikipedia through existing data and using simulations to explore potential…

计算与语言 · 计算机科学 2026-03-03 Siming Huang , Yuliang Xu , Mingmeng Geng , Yao Wan , Dongping Chen

On Wikipedia, sophisticated algorithmic tools are used to assess the quality of edits and take corrective actions. However, algorithms can fail to solve the problems they were designed for if they conflict with the values of communities who…

人机交互 · 计算机科学 2020-01-15 C. Estelle Smith , Bowen Yu , Anjali Srivastava , Aaron Halfaker , Loren Terveen , Haiyi Zhu

Large language models (LLMs) are trained on broad corpora and then used in communities with specialized norms. Is providing LLMs with community rules enough for models to follow these norms? We evaluate LLMs' capacity to detect (Task 1) and…

计算与语言 · 计算机科学 2026-05-11 Joshua Ashkinaze , Ruijia Guan , Laura Kurek , Eytan Adar , Ceren Budak , Eric Gilbert

Sentence level quality estimation (QE) for machine translation (MT) attempts to predict the translation edit rate (TER) cost of post-editing work required to correct MT output. We describe our view on sentence-level QE as dictated by…

计算与语言 · 计算机科学 2020-05-08 Junpei Zhou , Ciprian Chelba , Yuezhang , Li