中文
相关论文

相关论文: Modeling "Newsworthiness" for Lead-Generation Acro…

200 篇论文

Classic Topic Models are built under the Bag Of Words assumption, in which word position is ignored for simplicity. Besides, symmetric priors are typically used in most applications. In order to easily learn topics with different properties…

计算与语言 · 计算机科学 2018-06-27 Simón Roca-Sotelo , Jerónimo Arenas-García

Accuracy is one of the basic principles of journalism. However, it is increasingly hard to manage due to the diversity of news media. Some editors of online news tend to use catchy headlines which trick readers into clicking. These…

计算与语言 · 计算机科学 2017-08-30 Wei Wei , Xiaojun Wan

Among news disorders, propagandist news are particularly insidious, because they tend to mix oriented messages with factual reports intended to look like reliable news. To detect propaganda, extant approaches based on Language Models such…

Identifying risks associated with a company is important to investors and the well-being of the overall financial market. In this study, we build a computational framework to automatically extract company risk factors from news articles.…

计算与语言 · 计算机科学 2025-08-18 Jiaxin Pei , Soumya Vadlamannati , Liang-Kang Huang , Daniel Preotiuc-Pietro , Xinyu Hua

With the growth of the internet, the number of fake-news online has been proliferating every year. The consequences of such phenomena are manifold, ranging from lousy decision-making process to bullying and violence episodes. Therefore,…

信息检索 · 计算机科学 2018-09-05 Diego Esteves , Aniketh Janardhan Reddy , Piyush Chawla , Jens Lehmann

Most existing named entity recognition (NER) approaches are based on sequence labeling models, which focus on capturing the local context dependencies. However, the way of taking one sentence as input prevents the modeling of non-sequential…

计算与语言 · 计算机科学 2021-06-03 Zanbo Wang , Wei Wei , Xianling Mao , Shanshan Feng , Pan Zhou , Zhiyong He , Sheng Jiang

Increasing amounts of freely available data both in textual and relational form offers exploration of richer document representations, potentially improving the model performance and robustness. An emerging problem in the modern era is fake…

计算与语言 · 计算机科学 2022-02-16 Boshko Koloski , Timen Stepišnik-Perdih , Marko Robnik-Šikonja , Senja Pollak , Blaž Škrlj

Extracting coherent and human-understandable themes from large collections of unstructured historical newspaper archives presents significant challenges due to topic evolution, Optical Character Recognition (OCR) noise, and the sheer volume…

计算与语言 · 计算机科学 2025-12-15 Keerthana Murugaraj , Salima Lamsiyah , Marten During , Martin Theobald

We introduce a novel latent grouping model for predicting the relevance of a new document to a user. The model assumes a latent group structure for both users and documents. We compared the model against a state-of-the-art method, the User…

信息检索 · 计算机科学 2012-07-09 Eerika Savia , Kai Puolamaki , Janne Sinkkonen , Samuel Kaski

Issue tracking systems are used in the software industry for the facilitation of maintenance activities that keep the software robust and up to date with ever-changing industry requirements. Usually, users report issues that can be…

软件工程 · 计算机科学 2022-02-16 Anas Nadeem , Muhammad Usman Sarwar , Muhammad Zubair Malik

Most real-world document collections involve various types of metadata, such as author, source, and date, and yet the most commonly-used approaches to modeling text corpora ignore this information. While specialized models have been…

机器学习 · 统计学 2018-10-25 Dallas Card , Chenhao Tan , Noah A. Smith

While entity-oriented neural IR models have advanced significantly, they often overlook a key nuance: the varying degrees of influence individual entities within a document have on its overall relevance. Addressing this gap, we present…

信息检索 · 计算机科学 2024-01-12 Shubham Chatterjee , Iain Mackie , Jeff Dalton

This paper presents a modified neural model for topic detection from a corpus and proposes a new metric to evaluate the detected topics. The new model builds upon the embedded topic model incorporating some modifications such as document…

计算与语言 · 计算机科学 2023-06-09 Tomoya Kitano , Yuto Miyatake , Daisuke Furihata

In many domains such as medicine, training data is in short supply. In such cases, external knowledge is often helpful in building predictive models. We propose a novel method to incorporate publicly available domain expertise to build…

机器学习 · 计算机科学 2020-06-03 Yun Liu , Kun-Ta Chuang , Fu-Wen Liang , Huey-Jen Su , Collin M. Stultz , John V. Guttag

We consider a document classification problem where document labels are absent but only relevant keywords of a target class and unlabeled documents are given. Although heuristic methods based on pseudo-labeling have been considered,…

计算与语言 · 计算机科学 2019-10-31 Nontawat Charoenphakdee , Jongyeong Lee , Yiping Jin , Dittaya Wanvarie , Masashi Sugiyama

We study the problem of topic modeling in corpora whose documents are organized in a multi-level hierarchy. We explore a parametric approach to this problem, assuming that the number of topics is known or can be estimated by…

机器学习 · 统计学 2015-04-14 Do-kyum Kim , Geoffrey M. Voelker , Lawrence K. Saul

News Image Captioning aims to create captions from news articles and images, emphasizing the connection between textual context and visual elements. Recognizing the significance of human faces in news images and the face-name co-occurrence…

计算机视觉与模式识别 · 计算机科学 2024-03-22 Tingyu Qu , Tinne Tuytelaars , Marie-Francine Moens

This study proposes a Neural Attentive Bag-of-Entities model, which is a neural network model that performs text classification using entities in a knowledge base. Entities provide unambiguous and relevant semantic signals that are…

计算与语言 · 计算机科学 2019-09-11 Ikuya Yamada , Hiroyuki Shindo

Insightful findings in political science often require researchers to analyze documents of a certain subject or type, yet these documents are usually contained in large corpora that do not distinguish between pertinent and non-pertinent…

计算与语言 · 计算机科学 2019-10-29 Shrey Desai , Barea Sinno , Alex Rosenfeld , Junyi Jessy Li

Media houses reporting on public figures, often come with their own biases stemming from their respective worldviews. A characterization of these underlying patterns helps us in better understanding and interpreting news stories. For this,…

计算与语言 · 计算机科学 2023-09-13 Sharath Srivatsa , Srinath Srinivasa