中文
相关论文

相关论文: Modeling "Newsworthiness" for Lead-Generation Acro…

200 篇论文

Text classification methods have been widely investigated as a way to detect content of low credibility: fake news, social media bots, propaganda, etc. Quite accurate models (likely based on deep neural networks) help in moderating public…

计算与语言 · 计算机科学 2026-03-04 Piotr Przybyła , Alexander Shvets , Horacio Saggion

This paper presents a novel research problem on joint discovery of commonalities and differences between two individual documents (or document sets), called Comparative Document Analysis (CDA). Given any pair of documents from a document…

信息检索 · 计算机科学 2015-10-27 Xiang Ren , Yuanhua Lv , Kuansan Wang , Jiawei Han

Investigative journalists routinely confront large document collections. Large language models (LLMs) with retrieval-augmented generation (RAG) capabilities promise to accelerate the process of document discovery, but newsroom adoption…

信息检索 · 计算机科学 2025-10-01 Nick Hagar , Nicholas Diakopoulos , Jeremy Gilbert

Making disguise between real and fake news propagation through online social networks is an important issue in many applications. The time gap between the news release time and detection of its label is a significant step towards…

社会与信息网络 · 计算机科学 2019-09-06 Maryam Ramezani , Mina Rafiei , Soroush Omranpour , Hamid R. Rabiee

We consider the problem of modeling the content structure of texts within a specific domain, in terms of the topics the texts address and the order in which these topics appear. We first present an effective knowledge-lean method for…

计算与语言 · 计算机科学 2007-05-23 Regina Barzilay , Lillian Lee

Well curated, large-scale corpora of social media posts containing broad public opinion offer an alternative data source to complement traditional surveys. While surveys are effective at collecting representative samples and are capable of…

计算与语言 · 计算机科学 2025-02-14 Michael V. Arnold , Peter Sheridan Dodds , Christopher M. Danforth

Some news headlines mislead readers with overrated or false information, and identifying them in advance will better assist readers in choosing proper news stories to consume. This research introduces million-scale pairs of news headline…

计算与语言 · 计算机科学 2019-02-11 Seunghyun Yoon , Kunwoo Park , Joongbo Shin , Hongjun Lim , Seungpil Won , Meeyoung Cha , Kyomin Jung

News articles are driven by the informational sources journalists use in reporting. Modeling when, how and why sources get used together in stories can help us better understand the information we consume and even help journalists with the…

计算与语言 · 计算机科学 2023-05-25 Alexander Spangher , Nanyun Peng , Jonathan May , Emilio Ferrara

Lawyers and judges spend a large amount of time researching the proper legal authority to cite while drafting decisions. In this paper, we develop a citation recommendation tool that can help improve efficiency in the process of opinion…

信息检索 · 计算机科学 2021-06-22 Zihan Huang , Charles Low , Mengqiu Teng , Hongyi Zhang , Daniel E. Ho , Mark S. Krass , Matthias Grabmair

News headline generation is an essential problem of text summarization because it is constrained, well-defined, and is still hard to solve. Models with a limited vocabulary can not solve it well, as new named entities can appear regularly…

计算与语言 · 计算机科学 2020-05-06 Ilya Gusev

Debate portals and similar web platforms constitute one of the main text sources in computational argumentation research and its applications. While the corpora built upon these sources are rich of argumentatively relevant content and…

计算与语言 · 计算机科学 2020-11-04 Jonas Dorsch , Henning Wachsmuth

A crucial aspect of a rumor detection model is its ability to generalize, particularly its ability to detect emerging, previously unknown rumors. Past research has indicated that content-based (i.e., using solely source posts as input)…

计算与语言 · 计算机科学 2024-03-26 Yida Mu , Xingyi Song , Kalina Bontcheva , Nikolaos Aletras

We present a study on predicting the factuality of reporting and bias of news media. While previous work has focused on studying the veracity of claims or documents, here we are interested in characterizing entire news media. These are…

信息检索 · 计算机科学 2018-10-04 Ramy Baly , Georgi Karadzhov , Dimitar Alexandrov , James Glass , Preslav Nakov

Detecting and tracking emerging trends and weak signals in large, evolving text corpora is vital for applications such as monitoring scientific literature, managing brand reputation, surveilling critical infrastructure and more generally to…

计算与语言 · 计算机科学 2024-11-22 Allaa Boutaleb , Jerome Picault , Guillaume Grosjean

This paper proposes a new methodology to study sequential corpora by implementing a two-stage algorithm that learns time-based topics with respect to a scale of document positions and introduces the concept of Topic Scaling which ranks…

信息检索 · 计算机科学 2021-04-05 Sami Diaf , Ulrich Fritsche

News articles usually contain knowledge entities such as celebrities or organizations. Important entities in articles carry key messages and help to understand the content in a more direct way. An industrial news recommender system contains…

信息检索 · 计算机科学 2020-09-15 Danyang Liu , Jianxun Lian , Shiyin Wang , Ying Qiao , Jiun-Hung Chen , Guangzhong Sun , Xing Xie

In this work, we ask two questions: 1. Can we predict the type of community interested in a news article using only features from the article content? and 2. How well do these models generalize over time? To answer these questions, we…

信息检索 · 计算机科学 2018-08-29 Benjamin D. Horne , William Dron , Sibel Adali

Readability assessment aims to evaluate the reading difficulty of a text. In recent years, while deep learning technology has been gradually applied to readability assessment, most approaches fail to consider either the length of the text…

计算与语言 · 计算机科学 2025-11-27 Yurui Zheng , Yijun Chen , Shaohong Zhang

Legal documents are unstructured, use legal jargon, and have considerable length, making them difficult to process automatically via conventional text processing techniques. A legal document processing system would benefit substantially if…

Finding related published articles is an important task in any science, but with the explosion of new work in the biomedical domain it has become especially challenging. Most existing methodologies use text similarity metrics to identify…

信息检索 · 计算机科学 2016-11-07 Jesse M Lingeman , Hong Yu