English
Related papers

Related papers: Modeling "Newsworthiness" for Lead-Generation Acro…

200 papers

Text classification methods have been widely investigated as a way to detect content of low credibility: fake news, social media bots, propaganda, etc. Quite accurate models (likely based on deep neural networks) help in moderating public…

Computation and Language · Computer Science 2026-03-04 Piotr Przybyła , Alexander Shvets , Horacio Saggion

This paper presents a novel research problem on joint discovery of commonalities and differences between two individual documents (or document sets), called Comparative Document Analysis (CDA). Given any pair of documents from a document…

Information Retrieval · Computer Science 2015-10-27 Xiang Ren , Yuanhua Lv , Kuansan Wang , Jiawei Han

Investigative journalists routinely confront large document collections. Large language models (LLMs) with retrieval-augmented generation (RAG) capabilities promise to accelerate the process of document discovery, but newsroom adoption…

Information Retrieval · Computer Science 2025-10-01 Nick Hagar , Nicholas Diakopoulos , Jeremy Gilbert

Making disguise between real and fake news propagation through online social networks is an important issue in many applications. The time gap between the news release time and detection of its label is a significant step towards…

Social and Information Networks · Computer Science 2019-09-06 Maryam Ramezani , Mina Rafiei , Soroush Omranpour , Hamid R. Rabiee

We consider the problem of modeling the content structure of texts within a specific domain, in terms of the topics the texts address and the order in which these topics appear. We first present an effective knowledge-lean method for…

Computation and Language · Computer Science 2007-05-23 Regina Barzilay , Lillian Lee

Well curated, large-scale corpora of social media posts containing broad public opinion offer an alternative data source to complement traditional surveys. While surveys are effective at collecting representative samples and are capable of…

Computation and Language · Computer Science 2025-02-14 Michael V. Arnold , Peter Sheridan Dodds , Christopher M. Danforth

Some news headlines mislead readers with overrated or false information, and identifying them in advance will better assist readers in choosing proper news stories to consume. This research introduces million-scale pairs of news headline…

Computation and Language · Computer Science 2019-02-11 Seunghyun Yoon , Kunwoo Park , Joongbo Shin , Hongjun Lim , Seungpil Won , Meeyoung Cha , Kyomin Jung

News articles are driven by the informational sources journalists use in reporting. Modeling when, how and why sources get used together in stories can help us better understand the information we consume and even help journalists with the…

Computation and Language · Computer Science 2023-05-25 Alexander Spangher , Nanyun Peng , Jonathan May , Emilio Ferrara

Lawyers and judges spend a large amount of time researching the proper legal authority to cite while drafting decisions. In this paper, we develop a citation recommendation tool that can help improve efficiency in the process of opinion…

Information Retrieval · Computer Science 2021-06-22 Zihan Huang , Charles Low , Mengqiu Teng , Hongyi Zhang , Daniel E. Ho , Mark S. Krass , Matthias Grabmair

News headline generation is an essential problem of text summarization because it is constrained, well-defined, and is still hard to solve. Models with a limited vocabulary can not solve it well, as new named entities can appear regularly…

Computation and Language · Computer Science 2020-05-06 Ilya Gusev

Debate portals and similar web platforms constitute one of the main text sources in computational argumentation research and its applications. While the corpora built upon these sources are rich of argumentatively relevant content and…

Computation and Language · Computer Science 2020-11-04 Jonas Dorsch , Henning Wachsmuth

A crucial aspect of a rumor detection model is its ability to generalize, particularly its ability to detect emerging, previously unknown rumors. Past research has indicated that content-based (i.e., using solely source posts as input)…

Computation and Language · Computer Science 2024-03-26 Yida Mu , Xingyi Song , Kalina Bontcheva , Nikolaos Aletras

We present a study on predicting the factuality of reporting and bias of news media. While previous work has focused on studying the veracity of claims or documents, here we are interested in characterizing entire news media. These are…

Information Retrieval · Computer Science 2018-10-04 Ramy Baly , Georgi Karadzhov , Dimitar Alexandrov , James Glass , Preslav Nakov

Detecting and tracking emerging trends and weak signals in large, evolving text corpora is vital for applications such as monitoring scientific literature, managing brand reputation, surveilling critical infrastructure and more generally to…

Computation and Language · Computer Science 2024-11-22 Allaa Boutaleb , Jerome Picault , Guillaume Grosjean

This paper proposes a new methodology to study sequential corpora by implementing a two-stage algorithm that learns time-based topics with respect to a scale of document positions and introduces the concept of Topic Scaling which ranks…

Information Retrieval · Computer Science 2021-04-05 Sami Diaf , Ulrich Fritsche

News articles usually contain knowledge entities such as celebrities or organizations. Important entities in articles carry key messages and help to understand the content in a more direct way. An industrial news recommender system contains…

Information Retrieval · Computer Science 2020-09-15 Danyang Liu , Jianxun Lian , Shiyin Wang , Ying Qiao , Jiun-Hung Chen , Guangzhong Sun , Xing Xie

In this work, we ask two questions: 1. Can we predict the type of community interested in a news article using only features from the article content? and 2. How well do these models generalize over time? To answer these questions, we…

Information Retrieval · Computer Science 2018-08-29 Benjamin D. Horne , William Dron , Sibel Adali

Readability assessment aims to evaluate the reading difficulty of a text. In recent years, while deep learning technology has been gradually applied to readability assessment, most approaches fail to consider either the length of the text…

Computation and Language · Computer Science 2025-11-27 Yurui Zheng , Yijun Chen , Shaohong Zhang

Legal documents are unstructured, use legal jargon, and have considerable length, making them difficult to process automatically via conventional text processing techniques. A legal document processing system would benefit substantially if…

Computation and Language · Computer Science 2022-11-08 Vijit Malik , Rishabh Sanjay , Shouvik Kumar Guha , Angshuman Hazarika , Shubham Nigam , Arnab Bhattacharya , Ashutosh Modi

Finding related published articles is an important task in any science, but with the explosion of new work in the biomedical domain it has become especially challenging. Most existing methodologies use text similarity metrics to identify…

Information Retrieval · Computer Science 2016-11-07 Jesse M Lingeman , Hong Yu