中文
相关论文

相关论文: StRE: Self Attentive Edit Quality Prediction in Wi…

200 篇论文

Malicious sockpuppet detection on Wikipedia is critical to preserving access to reliable information on the internet and preventing the spread of disinformation. Prior machine learning approaches rely on stylistic and meta-data features,…

机器学习 · 计算机科学 2025-10-29 Luc Raszewski , Christine De Kock

Social media companies as well as authorities make extensive use of artificial intelligence (AI) tools to monitor postings of hate speech, celebrations of violence or profanity. Since AI software requires massive volumes of data to train…

计算与语言 · 计算机科学 2021-09-30 Hadeel Saadany , Constantin Orasan

Training models for the automatic correction of machine-translated text usually relies on data consisting of (source, MT, human post- edit) triplets providing, for each source sentence, examples of translation errors with the corresponding…

计算与语言 · 计算机科学 2018-03-21 Matteo Negri , Marco Turchi , Rajen Chatterjee , Nicola Bertoldi

We show that generating English Wikipedia articles can be approached as a multi- document summarization of source documents. We use extractive summarization to coarsely identify salient information and a neural abstractive model to generate…

计算与语言 · 计算机科学 2018-02-01 Peter J. Liu , Mohammad Saleh , Etienne Pot , Ben Goodrich , Ryan Sepassi , Lukasz Kaiser , Noam Shazeer

Discussion threads form a central part of the experience on many Web sites, including social networking sites such as Facebook and Google Plus and knowledge creation sites such as Wikipedia. To help users manage the challenge of allocating…

社会与信息网络 · 计算机科学 2013-04-18 Lars Backstrom , Jon Kleinberg , Lillian Lee , Cristian Danescu-Niculescu-Mizil

Wikipedia is a major source of information providing a large variety of content online, trusted by readers from around the world. Readers go to Wikipedia to get reliable information about different subjects, one of the most popular being…

社会与信息网络 · 计算机科学 2020-06-25 Pushkal Agarwal , Miriam Redi , Nishanth Sastry , Edward Wood , Andrew Blick

Wikidata is a multi-language knowledge base that is being edited and maintained by editors from different language communities. Due to the structured nature of its content, the contributions are in various forms, including manual edit,…

人机交互 · 计算机科学 2023-11-07 Jeffrey Jun-jie Ma , Charles Chuankai Zhang

User-generated, multi-paragraph writing is pervasive and important in many social media platforms (i.e. Amazon reviews, AirBnB host profiles, etc). Ensuring high-quality content is important. Unfortunately, content submitted by users is…

人机交互 · 计算机科学 2018-04-20 Hamed Nilforoshan , Eugene Wu

Aspect Sentiment Triplet Extraction (ASTE) is a new fine-grained sentiment analysis task that aims to extract triplets of aspect terms, sentiments, and opinion terms from review sentences. Recently, span-level models achieve gratifying…

计算与语言 · 计算机科学 2022-11-11 Yuqi Chen , Keming Chen , Xian Sun , Zequn Zhang

As entity type systems become richer and more fine-grained, we expect the number of types assigned to a given entity to increase. However, most fine-grained typing work has focused on datasets that exhibit a low degree of type multiplicity.…

计算与语言 · 计算机科学 2017-04-26 Maxim Rabinovich , Dan Klein

Automatic quality evaluation of Web information is a task with many fields of applications and of great relevance, especially in critical domains like the medical one. We move from the intuition that the quality of content of medical Web…

信息检索 · 计算机科学 2016-03-08 Vittoria Cozza , Marinella Petrocchi , Angelo Spognardi

Over the past 20 years, Wikipedia has gone from a rather outlandish idea to a major reference work, with more than 60 million articles across all languages, including nearly 7 million in English [Wiki01]. Around 27,000 of these articles…

历史与综述 · 数学 2024-12-31 David Eppstein , Joel Brewster Lewis , Russ Woodroofe , XOR'easter

Sarcasm is common in online discussions, yet difficult for machines to identify because the intended meaning often contradicts the literal wording. In this work, I study sarcasm detection using only classical machine learning methods and…

计算与语言 · 计算机科学 2026-01-26 Subrata Karmaker

Wikipedia, the Web's largest encyclopedia, frequently faces content disputes or malicious users seeking to subvert its integrity. Administrators can mitigate such disruptions by enforcing "page protection" that selectively limits…

计算机与社会 · 计算机科学 2023-10-20 Thorsten Ruprechter , Manoel Horta Ribeiro , Robert West , Denis Helic

Scientific publications are the primary means to communicate research discoveries, where the writing quality is of crucial importance. However, prior work studying the human editing process in this domain mainly focused on the abstract or…

计算与语言 · 计算机科学 2022-11-01 Chao Jiang , Wei Xu , Samuel Stevens

Online encyclopedia such as Wikipedia has become one of the best sources of knowledge. Much effort has been devoted to expanding and enriching the structured data by automatic information extraction from unstructured text in Wikipedia.…

信息检索 · 计算机科学 2014-06-26 Kezun Zhang , Yanghua Xiao , Hanghang Tong , Haixun Wang , Wei Wang

Multi-modal generative document parsing systems challenge traditional evaluation: unlike deterministic OCR or layout models, they often produce semantically correct yet structurally divergent outputs. Conventional metrics-CER, WER, IoU, or…

计算与语言 · 计算机科学 2025-09-25 Renyu Li , Antonio Jimeno Yepes , Yao You , Kamil Pluciński , Maximilian Operlejn , Crag Wolfe

Automatic postediting (APE) is an automated process to refine a given machine translation (MT). Recent findings present that existing APE systems are not good at handling high-quality MTs even for a language pair with abundant data…

计算与语言 · 计算机科学 2023-06-21 Baikjin Jung , Myungji Lee , Jong-Hyeok Lee , Yunsu Kim

We study the problem of entity salience by proposing the design and implementation of SWAT, a system that identifies the salient Wikipedia entities occurring in an input document. SWAT consists of several modules that are able to detect and…

信息检索 · 计算机科学 2019-05-17 Marco Ponza , Paolo Ferragina , Francesco Piccinno

Text editing, i.e., the process of modifying or manipulating text, is a crucial step in human writing process. In this paper, we study the problem of controlled text editing by natural language instruction. According to a given instruction…

计算与语言 · 计算机科学 2023-10-10 Xiang Chen , Zheng Li , Xiaojun Wan