中文
相关论文

相关论文: WikiSQE: A Large-Scale Dataset for Sentence Qualit…

200 篇论文

Huge numbers of new words emerge every day, leading to a great need for representing them with semantic meaning that is understandable to NLP systems. Sememes are defined as the minimum semantic units of human languages, the combination of…

计算与语言 · 计算机科学 2018-08-17 Wei Li , Xuancheng Ren , Damai Dai , Yunfang Wu , Houfeng Wang , Xu Sun

Content moderation in online platforms is crucial for ensuring activity therein adheres to existing policies, especially as these platforms grow. NLP research in this area has typically focused on automating some part of it given that it is…

计算与语言 · 计算机科学 2024-08-13 Hsuvas Borkakoty , Luis Espinosa-Anke

Understanding whether a generated table is of good quality is important to be able to use it in creating or editing documents using automatic methods. In this work, we underline that existing measures for table quality evaluation fail to…

计算与语言 · 计算机科学 2024-11-26 Pritika Ramu , Aparna Garimella , Sambaran Bandyopadhyay

The quality of a document is affected by various factors, including grammaticality, readability, stylistics, and expertise depth, making the task of document quality assessment a complex one. In this paper, we explore this task in the…

计算与语言 · 计算机科学 2019-01-15 Aili Shen , Bahar Salehi , Timothy Baldwin , Jianzhong Qi

Wikipedia serves as a good example of how editors collaborate to form and maintain an article. The relationship between editors, derived from their sequence of editing activity, results in a directed network structure called the revision…

社会与信息网络 · 计算机科学 2019-04-18 James R. Ashford , Liam D. Turner , Roger M. Whitaker , Alun Preece , Diane Felmlee , Don Towsley

One of the biggest challenges hindering progress in low-resource and multilingual machine translation is the lack of good evaluation benchmarks. Current evaluation benchmarks either lack good coverage of low-resource languages, consider…

With so many articles of varying qualities being produced every moment, it is a very urgent task to screen outstanding articles and commit them to social media. To our best knowledge, there is a lack of datasets and mature research works in…

计算与语言 · 计算机科学 2022-08-25 Chunhui Ai , Derui Wang , Xu Yan , Yang Xu , Wenrui Xie , Ziqiang Cao

Wikipedia is a popular web-based encyclopedia edited freely and collaboratively by its users. In this paper we present an analysis of Wikipedias in several languages as complex networks. The hyperlinks pointing from one Wikipedia article to…

物理与社会 · 物理学 2009-11-11 V. Zlatic , M. Bozicevic , H. Stefancic , M. Domazet

Collaborative content creation inevitably reaches situations where different points of view lead to conflict. We focus on Wikipedia, the free encyclopedia anyone may edit, where disputes about content in controversial articles often reflect…

NLP researchers need more, higher-quality text datasets. Human-labeled datasets are expensive to collect, while datasets collected via automatic retrieval from the web such as WikiBio are noisy and can include undesired biases. Moreover,…

计算与语言 · 计算机科学 2022-01-14 Ann Yuan , Daphne Ippolito , Vitaly Nikolaev , Chris Callison-Burch , Andy Coenen , Sebastian Gehrmann

Wikipedia is a major source of information providing a large variety of content online, trusted by readers from around the world. Readers go to Wikipedia to get reliable information about different subjects, one of the most popular being…

社会与信息网络 · 计算机科学 2020-06-25 Pushkal Agarwal , Miriam Redi , Nishanth Sastry , Edward Wood , Andrew Blick

We perform an in-depth analysis on the inequality in 863 Wikimedia projects. We take the complete editing history of 267,304,095 Wikimedia items until 2016, which not only covers every language edition of Wikipedia, but also embraces the…

物理与社会 · 物理学 2018-12-18 Jinhyuk Yun , Sang Hoon Lee , Hawoong Jeong

Peer review is a cornerstone of quality control in scientific publishing. With the increasing workload, the unintended use of `quick' heuristics, referred to as lazy thinking, has emerged as a recurring issue compromising review quality.…

计算与语言 · 计算机科学 2025-06-04 Sukannya Purkayastha , Zhuang Li , Anne Lauscher , Lizhen Qu , Iryna Gurevych

Text simplification systems generate versions of texts that are easier to understand for a broader audience. The quality of simplified texts is generally estimated using metrics that compare to human references, which can be difficult to…

计算与语言 · 计算机科学 2020-12-24 Reno Kriz , Marianna Apidianaki , Chris Callison-Burch

We propose a method to determine whether a given article was written entirely by a generative language model or perhaps contains edits by a different author, possibly a human. Our process involves multiple tests for the origin of individual…

信息论 · 计算机科学 2024-08-27 Idan Kashtan , Alon Kipnis

The Internet-based encyclopaedia Wikipedia has grown to become one of the most visited web-sites on the Internet. However, critics have questioned the quality of entries, and an empirical study has shown Wikipedia to contain errors in a…

数字图书馆 · 计算机科学 2011-01-04 Finn Aarup Nielsen

Knowledge Editing (KE) aims to adjust a Large Language Model's (LLM) internal representations and parameters to correct inaccuracies and improve output consistency without incurring the computational expense of re-training the entire model.…

计算与语言 · 计算机科学 2025-05-29 Liyu Zhang , Weiqi Wang , Tianqing Fang , Yangqiu Song

We introduce and formalize the Synthetic Dataset Quality Estimation (SynQuE) problem: ranking synthetic datasets by their expected real-world task performance using only limited unannotated real data. This addresses a critical and open…

机器学习 · 计算机科学 2026-05-04 Arthur Chen , Victor Zhong

We propose an automatic language-independent graph-based method to build \`a-la-carte article collections on user-defined domains from the Wikipedia. The core model is based on the exploration of the encyclopaedia's category graph and can…

计算与语言 · 计算机科学 2020-05-05 Cristina España-Bonet , Alberto Barrón-Cedeño , Lluís Màrquez

This paper investigates two complementary paradigms for predicting machine translation (MT) quality: source-side difficulty prediction and candidate-side quality estimation (QE). The rapid adoption of Large Language Models (LLMs) into MT…

计算与语言 · 计算机科学 2026-03-05 Malik Marmonier , Benoît Sagot , Rachel Bawden