中文
相关论文

相关论文: NewsEdits: A Dataset of Revision Histories for New…

200 篇论文

News article revision histories provide clues to narrative and factual evolution in news articles. To facilitate analysis of this evolution, we present the first publicly available dataset of news revision histories, NewsEdits. Our dataset…

计算与语言 · 计算机科学 2022-06-16 Alexander Spangher , Xiang Ren , Jonathan May , Nanyun Peng

We release a corpus of 43 million atomic edits across 8 languages. These edits are mined from Wikipedia edit history and consist of instances in which a human editor has inserted a single contiguous phrase into, or deleted a single…

计算与语言 · 计算机科学 2018-08-29 Manaal Faruqui , Ellie Pavlick , Ian Tenney , Dipanjan Das

Understanding the writing frame of news articles is vital for addressing social issues, and thus has attracted notable attention in the fields of communication studies. Yet, assessing such news article frames remains a challenge due to the…

计算与语言 · 计算机科学 2024-05-24 Xi Chen , Mattia Samory , Scott Hale , David Jurgens , Przemyslaw A. Grabowicz

We present NEWSROOM, a summarization dataset of 1.3 million articles and summaries written by authors and editors in newsrooms of 38 major news publications. Extracted from search and social media metadata between 1998 and 2017, these…

计算与语言 · 计算机科学 2020-05-19 Max Grusky , Mor Naaman , Yoav Artzi

As events progress, news articles often update with new information: if we are not cautious, we risk propagating outdated facts. In this work, we hypothesize that linguistic features indicate factual fluidity, and that we can predict which…

计算与语言 · 计算机科学 2024-12-02 Alexander Spangher , Kung-Hsiang Huang , Hyundong Cho , Jonathan May

To foster the development of new models for collaborative AI-assisted report generation, we introduce MegaWika, consisting of 13 million Wikipedia articles in 50 diverse languages, along with their 71 million referenced source materials. We…

We present a dataset that contains every instance of all tokens (~ words) ever written in undeleted, non-redirect English Wikipedia articles until October 2016, in total 13,545,349,787 instances. Each token is annotated with (i) the article…

计算与语言 · 计算机科学 2017-03-27 Fabian Flöck , Kenan Erdogan , Maribel Acosta

Writing a scientific article is a challenging task as it is a highly codified and specific genre, consequently proficiency in written communication is essential for effectively conveying research findings and ideas. In this article, we…

计算与语言 · 计算机科学 2025-01-10 Leane Jourdan , Florian Boudin , Nicolas Hernandez , Richard Dufour

We develop novel annotation guidelines for sentence-level subjectivity detection, which are not limited to language-specific cues. We use our guidelines to collect NewsSD-ENG, a corpus of 638 objective and 411 subjective sentences extracted…

Scientific publications are the primary means to communicate research discoveries, where the writing quality is of crucial importance. However, prior work studying the human editing process in this domain mainly focused on the abstract or…

计算与语言 · 计算机科学 2022-11-01 Chao Jiang , Wei Xu , Samuel Stevens

In the U.S. historically, local newspapers drew their content largely from newswires like the Associated Press. Historians argue that newswires played a pivotal role in creating a national identity and shared understanding of the world, but…

计算与语言 · 计算机科学 2024-06-17 Emily Silcock , Abhishek Arora , Luca D'Amico-Wong , Melissa Dell

We present the Verifee Dataset: a novel dataset of news articles with fine-grained trustworthiness annotations. We develop a detailed methodology that assesses the texts based on their parameters encompassing editorial transparency,…

计算与语言 · 计算机科学 2022-12-19 Matyáš Boháček , Michal Bravanský , Filip Trhlík , Václav Moravec

To improve software engineering, software repositories have been mined for code snippets and bug fixes. Typically, this mining takes place at the level of files or commits. To be able to dig deeper and to extract insights at a higher…

软件工程 · 计算机科学 2020-05-07 Sebastian Baltes , Markus Wagner

In this paper, we introduce a dataset of multilingual news articles covering the 2021 Tokyo Olympics. A total of 10,940 news articles were gathered from 1,918 different publishers, covering 1,350 sub-events of the 2021 Olympics, and…

信息检索 · 计算机科学 2025-02-17 Erik Novak , Erik Calcina , Dunja Mladenić , Marko Grobelnik

Despite increasing awareness and research around fake news, there is still a significant need for datasets that specifically target racial slurs and biases within North American political speeches. This is particulary important in the…

计算与语言 · 计算机科学 2024-01-09 Shaina Raza , Mizanur Rahman , Shardul Ghuge

In order to simplify a sentence, human editors perform multiple rewriting transformations: they split it into several shorter sentences, paraphrase words (i.e. replacing complex words or phrases by simpler synonyms), reorder components,…

计算与语言 · 计算机科学 2020-05-04 Fernando Alva-Manchego , Louis Martin , Antoine Bordes , Carolina Scarton , Benoît Sagot , Lucia Specia

In this paper, we present a dataset of 713k articles collected between 02/2018-11/2018. These articles are collected directly from 194 news and media outlets including mainstream, hyper-partisan, and conspiracy sources. We incorporate…

计算机与社会 · 计算机科学 2019-04-03 Jeppe Norregaard , Benjamin D. Horne , Sibel Adali

A diversity of tasks use language models trained on semantic similarity data. While there are a variety of datasets that capture semantic similarity, they are either constructed from modern web data or are relatively small datasets created…

计算与语言 · 计算机科学 2023-08-25 Emily Silcock , Melissa Dell

This paper describes a novel dataset consisting of sentences with semantic similarity annotations. The data originate from the journalistic domain in the Czech language. We describe the process of collecting and annotating the data in…

计算与语言 · 计算机科学 2022-01-24 Jakub Sido , Michal Seják , Ondřej Pražák , Miloslav Konopík , Václav Moravec

Though exponentially growing health-related literature has been made available to a broad audience online, the language of scientific articles can be difficult for the general public to understand. Therefore, adapting this expert-level…

计算与语言 · 计算机科学 2022-10-25 Kush Attal , Brian Ondov , Dina Demner-Fushman
‹ 上一页 1 2 3 10 下一页 ›