English
Related papers

Related papers: Microblog Topic Identification using Linked Open D…

200 papers

Urban transit agencies increasingly turn to social media to monitor emerging service risks such as crowding, delays, and safety incidents, yet the signals of concern are sparse, short, and easily drowned by routine chatter. We address this…

Machine Learning · Computer Science 2025-12-09 Fatima Ashraf , Muhammad Ayub Sabir , Jiaxin Deng , Junbiao Pang , Haitao Yu

People are shifting from traditional news sources to online news at an incredibly fast rate. However, the technology behind online news consumption promotes content that confirms the users' existing point of view. This phenomenon has led to…

Social and Information Networks · Computer Science 2017-11-29 Preethi Lahoti , Kiran Garimella , Aristides Gionis

This article presents a novel approach for learning low-dimensional distributed representations of users in online social networks. Existing methods rely on the network structure formed by the social relationships among users to extract…

Social and Information Networks · Computer Science 2017-10-23 Harvineet Singh , Amitabha Bagchi , Parag Singla

When dealing with large collections of documents, it is imperative to quickly get an overview of the texts' contents. In this paper we show how this can be achieved by using a clustering algorithm to identify topics in the dataset and then…

Computation and Language · Computer Science 2017-07-20 Franziska Horn , Leila Arras , Grégoire Montavon , Klaus-Robert Müller , Wojciech Samek

We propose a straightforward solution for detecting scarce topics in unbalanced short-text datasets. Our approach, named CWUTM (Topic model based on co-occurrence word networks for unbalanced short text datasets), Our approach addresses the…

Computation and Language · Computer Science 2023-11-07 Chengjie Ma , Junping Du , Meiyu Liang , Zeli Guan

The task of discovering topics in text corpora has been dominated by Latent Dirichlet Allocation and other Topic Models for over a decade. In order to apply these approaches to massive text corpora, the vocabulary needs to be reduced…

Computation and Language · Computer Science 2019-08-08 Gibran Fuentes-Pineda , Ivan Vladimir Meza-Ruiz

As the emergence and the thriving development of social networks, a huge number of short texts are accumulated and need to be processed. Inferring latent topics of collected short texts is useful for understanding its hidden structure and…

Machine Learning · Statistics 2018-04-04 Zhenghang Cui , Issei Sato , Masashi Sugiyama

In this paper, we propose a framework to infer the topic preferences of Donald Trump's followers on Twitter. We first use latent Dirichlet allocation (LDA) to derive the weighted mixture of topics for each Trump tweet. Then we use negative…

Social and Information Networks · Computer Science 2016-03-11 Yu Wang , Jiebo Luo , Richard Niemi , Yuncheng Li , Tianran Hu

Social-media data provides increasing opportunities for automated analysis of large sets of textual documents. So far, automated tools have been developed to account for either the social networks between the participants of the debates, or…

Digital Libraries · Computer Science 2017-11-23 Iina Hellsten , Loet Leydesdorff

City Logistics is characterized by multiple stakeholders that often have different views of such a complex system. From a public policy perspective, identifying stakeholders, issues and trends is a daunting challenge, only partially…

Machine Learning · Computer Science 2019-06-19 Simon Tamayo , François Combes , Gaudron Arthur

In the real world, many topics are inter-correlated, making it challenging to investigate their structure and relationships. Understanding the interplay between topics and their relevance can provide valuable insights for researchers,…

Applications · Statistics 2024-02-01 Yeseul Jeon , Jina Park , Ick Hoon Jin , Dongjun Chungc

Social media platforms contain a great wealth of information which provides opportunities for us to explore hidden patterns or unknown correlations, and understand people's satisfaction with what they are discussing. As one showcase, in…

Information Retrieval · Computer Science 2017-05-24 Zhengkui Wang , Guangdong Bai , Soumyadeb Chowdhury , Quanqing Xu , Zhi Lin Seow

Computer users are generally faced with difficulties in making correct security decisions. While an increasingly fewer number of people are trying or willing to take formal security training, online sources including news, security blogs,…

Cryptography and Security · Computer Science 2020-06-29 Tingmin Wu , Wanlun Ma , Sheng Wen , Xin Xia , Cecile Paris , Surya Nepal , Yang Xiang

Language Identification (LID) is a challenging task, especially when the input texts are short and noisy such as posts and statuses on social media or chat logs on gaming forums. The task has been tackled by either designing a feature set…

Computation and Language · Computer Science 2019-10-16 Duy Tin Vo , Richard Khoury

Many data sets contain rich information about objects, as well as pairwise relations between them. For instance, in networks of websites, scientific papers, and other documents, each node has content consisting of a collection of words, as…

Machine Learning · Computer Science 2014-10-30 Yaojia Zhu , Xiaoran Yan , Lise Getoor , Cristopher Moore

Twitter messages often contain so-called hashtags to denote keywords related to them. Using a dataset of 29 million messages, I explore relations among these hashtags with respect to co-occurrences. Furthermore, I present an attempt to…

Computation and Language · Computer Science 2011-11-29 Jan Pöschko

Social media datasets, especially Twitter tweets, are popular in the field of text classification. Tweets are a valuable source of micro-text (sometimes referred to as "micro-blogs"), and have been studied in domains such as sentiment…

Information Retrieval · Computer Science 2017-08-29 Ankit Vadehra , Maura R. Grossman , Gordon V. Cormack

Probabilistic models can learn users' preferences from the history of their item adoptions on a social media site, and in turn, recommend new items to users based on learned preferences. However, current models ignore psychological factors…

Information Retrieval · Computer Science 2013-11-07 Jeon-Hyung Kang , Kristina Lerman

This paper proposes a topic modeling method that scales linearly to billions of documents. We make three core contributions: i) we present a topic modeling method, Tensor Latent Dirichlet Allocation (TLDA), that has identifiable and…

Machine Learning · Computer Science 2026-01-14 Sara Kangaslahti , Danny Ebanks , Jean Kossaifi , Anqi Liu , R. Michael Alvarez , Animashree Anandkumar

Academic researchers often need to face with a large collection of research papers in the literature. This problem may be even worse for postgraduate students who are new to a field and may not know where to start. To address this problem,…

Computation and Language · Computer Science 2016-09-30 Leonard K. M. Poon , Nevin L. Zhang