English
Related papers

Related papers: Using machine learning and information visualisati…

200 papers

Large language models (LLMs) offer new opportunities for scalable analysis of online discourse. Yet their use in multilingual social science research remains constrained by model size, cost and linguistic bias. We develop a lightweight,…

Computation and Language · Computer Science 2025-12-30 Andrea Nasuto , Stefano Maria Iacus , Francisco Rowe , Devika Jain

Content polluters, or bots that hijack a conversation for political or advertising purposes are a known problem for event prediction, election forecasting and when distinguishing real news from fake news in social media data. Identifying…

Social and Information Networks · Computer Science 2018-04-18 Mehwish Nasim , Andrew Nguyen , Nick Lothian , Robert Cope , Lewis Mitchell

Over the last decade, similar to other application domains, social media content has been proven very effective in disaster informatics. However, due to the unstructured nature of the data, several challenges are associated with disaster…

Computation and Language · Computer Science 2024-05-03 Ayaz Mehmood , Muhammad Tayyab Zamir , Muhammad Asif Ayub , Nasir Ahmad , Kashif Ahmad

Topic models, such as latent Dirichlet allocation (LDA), can be useful tools for the statistical analysis of document collections and other discrete data. The LDA model assumes that the words of each document arise from a mixture of topics,…

Applications · Statistics 2009-09-29 David M. Blei , John D. Lafferty

Urban transit agencies increasingly turn to social media to monitor emerging service risks such as crowding, delays, and safety incidents, yet the signals of concern are sparse, short, and easily drowned by routine chatter. We address this…

Machine Learning · Computer Science 2025-12-09 Fatima Ashraf , Muhammad Ayub Sabir , Jiaxin Deng , Junbiao Pang , Haitao Yu

According to tastes, a person could show preference for a given category of content to a greater or lesser extent. However, quantifying people's amount of interest in a certain topic is a challenging task, especially considering the massive…

Social and Information Networks · Computer Science 2018-07-20 Lorena Recalde , Ricardo Baeza-Yates

Aspect-based opinion mining is widely applied to review data to aggregate or summarize opinions of a product, and the current state-of-the-art is achieved with Latent Dirichlet Allocation (LDA)-based model. Although social media data like…

Computation and Language · Computer Science 2016-09-22 Kar Wai Lim , Wray Buntine

Opinion prediction on Twitter is challenging due to the transient nature of tweet content and neighbourhood context. In this paper, we model users' tweet posting behaviour as a temporal point process to jointly predict the posting time and…

Social and Information Networks · Computer Science 2020-05-28 Lixing Zhu , Yulan He , Deyu Zhou

Twitter is one of the most popular microblogging services in the world. The great amount of information within Twitter makes it an important information channel for people to learn and share news. Twitter hashtag is an popular feature that…

Social and Information Networks · Computer Science 2018-05-01 Shih-Feng Yang , Julia Taylor Rayz

Financial news items are unstructured sources of information that can be mined to extract knowledge for market screening applications. Manual extraction of relevant information from the continuous stream of finance-related news is…

This article charts the work of a 4 month project aimed at automatically identifying patterns of tweets popularity evolution using Machine Learning and Deep Learning techniques. To apprehend both the data and the extent of the problem, a…

Machine Learning · Computer Science 2023-01-04 Ferdinand Willemin

The digital town hall of Twitter becomes a preferred medium of communication for individuals and organizations across the globe. Some of them reach audiences of millions, while others struggle to get noticed. Given the impact of social…

Social and Information Networks · Computer Science 2021-02-23 Damian Konrad Kowalczyk , Jan Larsen

Traditionally, Latent Dirichlet Allocation (LDA) ingests words in a collection of documents to discover their latent topics using word-document co-occurrences. However, it is unclear how to achieve the best results for languages without…

Computation and Language · Computer Science 2021-08-25 Jin Cheevaprawatdomrong , Alexandra Schofield , Attapol T. Rutherford

In Twitter, and other microblogging services, the generation of new content by the crowd is often biased towards immediacy: what is happening now. Prompted by the propagation of commentary and information through multiple mediums, users on…

Information Retrieval · Computer Science 2016-02-10 Flávio Martins , João Magalhães , Jamie Callan

Public concern detection provides potential guidance to the authorities for crisis management before or during a pandemic outbreak. Detecting people's concerns and attention from online social media platforms has been widely acknowledged as…

Computation and Language · Computer Science 2021-06-21 Jingli Shi , Weihua Li , Sira Yongchareon , Yi Yang , Quan Bai

Twitter has become one of the most sought after places to discuss a wide variety of topics, including medically relevant issues such as cancer. This helps spread awareness regarding the various causes, cures and prevention methods of…

Social and Information Networks · Computer Science 2021-01-27 Rakesh Bal , Sayan Sinha , Swastika Dutta , Rishabh Joshi , Sayan Ghosh , Ritam Dutt

This paper introduces a large collection of time series data derived from Twitter, postprocessed using word embedding techniques, as well as specialized fine-tuned language models. This data comprises the past five years and captures…

Computation and Language · Computer Science 2023-08-07 Daniel Loureiro , Kiamehr Rezaee , Talayeh Riahi , Francesco Barbieri , Leonardo Neves , Luis Espinosa Anke , Jose Camacho-Collados

In this paper , we tackle Sentiment Analysis conditioned on a Topic in Twitter data using Deep Learning . We propose a 2-tier approach : In the first phase we create our own Word Embeddings and see that they do perform better than…

Computation and Language · Computer Science 2017-10-31 Sharath T. S. , Shubhangi Tandon

The legislative output of Colombia's House of Representatives between 2014 and 2025 is analyzed using 4,083 bills. Bipartite networks are constructed between parties and bills, and between representatives and bills, along with their…

Physics and Society · Physics 2025-12-19 Juan Sosa , Brayan Riveros , Emma J. Camargo-Díaz

During the 2016 US elections Twitter experienced unprecedented levels of propaganda and fake news through the collaboration of bots and hired persons, the ramifications of which are still being debated. This work proposes an approach to…

Social and Information Networks · Computer Science 2017-11-30 Erdem Beğenilmiş , Suzan Üsküdarlı