中文
相关论文

相关论文: A Massive Scale Semantic Similarity Dataset of His…

200 篇论文

In the U.S. historically, local newspapers drew their content largely from newswires like the Associated Press. Historians argue that newswires played a pivotal role in creating a national identity and shared understanding of the world, but…

计算与语言 · 计算机科学 2024-06-17 Emily Silcock , Abhishek Arora , Luca D'Amico-Wong , Melissa Dell

Existing full text datasets of U.S. public domain newspapers do not recognize the often complex layouts of newspaper scans, and as a result the digitized content scrambles texts from articles, headlines, captions, advertisements, and other…

Understanding the writing frame of news articles is vital for addressing social issues, and thus has attracted notable attention in the fields of communication studies. Yet, assessing such news article frames remains a challenge due to the…

计算与语言 · 计算机科学 2024-05-24 Xi Chen , Mattia Samory , Scott Hale , David Jurgens , Przemyslaw A. Grabowicz

This paper describes a novel dataset consisting of sentences with semantic similarity annotations. The data originate from the journalistic domain in the Czech language. We describe the process of collecting and annotating the data in…

计算与语言 · 计算机科学 2022-01-24 Jakub Sido , Michal Seják , Ondřej Pražák , Miloslav Konopík , Václav Moravec

The degree of semantic relatedness of two units of language has long been considered fundamental to understanding meaning. Additionally, automatically determining relatedness has many applications such as question answering and…

计算与语言 · 计算机科学 2023-03-21 Mohamed Abdalla , Krishnapriya Vishnubhotla , Saif M. Mohammad

Measuring the congruence between two texts has several useful applications, such as detecting the prevalent deceptive and misleading news headlines on the web. Many works have proposed machine learning based solutions such as text…

计算与语言 · 计算机科学 2020-10-09 Rahul Mishra , Piyush Yadav , Remi Calizzano , Markus Leippold

We find that existing language modeling datasets contain many near-duplicate examples and long repetitive substrings. As a result, over 1% of the unprompted output of language models trained on these datasets is copied verbatim from the…

Millions of news articles are published online every day, which can be overwhelming for readers to follow. Grouping articles that are reporting the same event into news stories is a common way of assisting readers in their news consumption.…

计算与语言 · 计算机科学 2020-04-15 Xiaotao Gu , Yuning Mao , Jiawei Han , Jialu Liu , Hongkun Yu , You Wu , Cong Yu , Daniel Finnie , Jiaqi Zhai , Nicholas Zukoski

This paper investigates sentence-level text reuse in multilingual journalism, analyzing where reused content occurs within articles. We present a weakly supervised method for detecting sentence-level cross-lingual reuse without requiring…

计算与语言 · 计算机科学 2026-04-01 Soveatin Kuntur , Nina Smirnova , Anna Wroblewska , Philipp Mayr , Sebastijan Razboršek Maček

Newspapers are a popular form of written discourse, read by many people, thanks to the novelty of the information provided by the news content in it. A headline is the most widely read part of any newspaper due to its appearance in a bigger…

计算与语言 · 计算机科学 2019-10-21 Elizabeth Jasmi George , Radhika Mamidi

Semantic sentence embeddings are usually supervisedly built minimizing distances between pairs of embeddings of sentences labelled as semantically similar by annotators. Since big labelled datasets are rare, in particular for non-English…

计算与语言 · 计算机科学 2021-10-06 Marco Di Giovanni , Marco Brambilla

Timeline generation is of great significance for a comprehensive understanding of the development of events over time. Its goal is to organize news chronologically, which helps to identify patterns and trends that may be obscured when…

信息检索 · 计算机科学 2025-02-12 Xiaochen Liu , Yanan Zhang

We propose a novel method for generating titles for unstructured text documents. We reframe the problem as a sequential question-answering task. A deep neural network is trained on document-title pairs with decomposable titles, meaning that…

计算与语言 · 计算机科学 2019-05-13 Oleg Vasilyev , Tom Grek , John Bohannon

News articles capture a variety of topics about our society. They reflect not only the socioeconomic activities that happened in our physical world, but also some of the cultures, human interests, and public concerns that exist only in the…

社会与信息网络 · 计算机科学 2018-09-11 Yingjie Hu , Xinyue Ye , Shih-Lung Shaw

News websites make editorial decisions about what stories to include on their website homepages and what stories to emphasize (e.g., large font size for main story). The emphasized stories on a news website are often highly similar to many…

数字图书馆 · 计算机科学 2018-07-03 Grant C. Atkins , Alexander Nwala , Michele C. Weigle , Michael L. Nelson

Detecting implicit causal relations in texts is a task that requires both common sense and world knowledge. Existing datasets are focused either on commonsense causal reasoning or explicit causal relations. In this work, we present…

计算与语言 · 计算机科学 2021-09-29 Ilya Gusev , Alexey Tikhonov

Chronicling America is a product of the National Digital Newspaper Program, a partnership between the Library of Congress and the National Endowment for the Humanities to digitize historic newspapers. Over 16 million pages of historic…

Question answering (QA) and Machine Reading Comprehension (MRC) tasks have significantly advanced in recent years due to the rapid development of deep learning techniques and, more recently, large language models. At the same time, many…

计算与语言 · 计算机科学 2024-05-13 Bhawna Piryani , Jamshid Mozafari , Adam Jatowt

English news headlines form a register with unique syntactic properties that have been documented in linguistics literature since the 1930s. However, headlines have received surprisingly little attention from the NLP syntactic parsing…

计算与语言 · 计算机科学 2023-01-26 Adrian Benton , Tianze Shi , Ozan İrsoy , Igor Malioutov

We present MediaSpin, a large-scale language resource capturing how major news outlets modify headlines after publication, and MediaSpin-in-the-Wild, a complementary dataset linking these revised headlines to their downstream engagement on…

计算与语言 · 计算机科学 2026-05-18 Preetika Verma , Kokil Jaidka
‹ 上一页 1 2 3 10 下一页 ›