中文
相关论文

相关论文: Short Text Topic Modeling: Application to tweets a…

200 篇论文

Twitter, a microblogging service, is todays most popular platform for communication in the form of short text messages, called Tweets. Users use Twitter to publish their content either for expressing concerns on information news or views on…

社会与信息网络 · 计算机科学 2017-11-29 Dhanasekar Sundararaman , Priya Arora , Vishwanath Seshagiri

Topic modeling refers to the task of discovering the underlying thematic structure in a text corpus, where the output is commonly presented as a report of the top terms appearing in each topic. Despite the diversity of topic modeling…

机器学习 · 计算机科学 2014-06-20 Derek Greene , Derek O'Callaghan , Pádraig Cunningham

Extracting topics from text has become an essential task, especially with the rapid growth of unstructured textual data. Most existing works rely on highly computational methods to address this challenge. In this paper, we argue that…

计算与语言 · 计算机科学 2025-11-07 Salma Mekaoui , Hiba Sofyan , Imane Amaaz , Imane Benchrif , Arsalane Zarghili , Ilham Chaker , Nikola S. Nikolov

Twitter data is extremely noisy -- each tweet is short, unstructured and with informal language, a challenge for current topic modeling. On the other hand, tweets are accompanied by extra information such as authorship, hashtags and the…

计算与语言 · 计算机科学 2016-09-23 Kar Wai Lim , Changyou Chen , Wray Buntine

Bitcoin has increased investment interests in people during the last decade. We have seen an increase in the number of posts on social media platforms about cryptocurrency, especially Bitcoin. This project focuses on analyzing user tweet…

人工智能 · 计算机科学 2024-12-04 Ashutosh Hathidara , Gaurav Atavale , Suyash Chaudhary

Financial analyses of stock markets rely heavily on quantitative approaches in an attempt to predict subsequent or market movements based on historical prices and other measurable metrics. These quantitative analyses might have missed out…

计算与语言 · 计算机科学 2020-08-04 Shaan Aryaman , Nguwi Yok Yen

Topic models are popular statistical tools for detecting latent semantic topics in a text corpus. They have been utilized in various applications across different fields. However, traditional topic models have some limitations, including…

计算与语言 · 计算机科学 2023-10-10 Pritom Saha Akash , Trisha Das , Kevin Chen-Chuan Chang

Conventional topic models are ineffective for topic extraction from microblog messages, because the data sparseness exhibited in short messages lacking structure and contexts results in poor message-level word co-occurrence patterns. To…

计算与语言 · 计算机科学 2018-09-12 Jing Li , Yan Song , Zhongyu Wei , Kam-Fai Wong

Social media platforms such as Twitter (now X) provide rich data for analyzing public discourse, especially during crises such as the COVID-19 pandemic. However, the brevity, informality, and noise of social media short texts often hinder…

计算与语言 · 计算机科学 2025-10-23 Wangjiaxuan Xin , Shuhua Yin , Shi Chen , Yaorong Ge

This study investigates the relationship between narratives conveyed through microblogging platforms, namely Twitter, and the value of crypto assets. Our study provides a unique technique to build narratives about cryptocurrency by…

计算金融 · 定量金融 2023-06-12 Lubdhak Mondal , Udeshya Raj , Abinandhan S , Began Gowsik S , Sarwesh P , Abhijeet Chandra

Many data sets contain rich information about objects, as well as pairwise relations between them. For instance, in networks of websites, scientific papers, and other documents, each node has content consisting of a collection of words, as…

机器学习 · 计算机科学 2014-10-30 Yaojia Zhu , Xiaoran Yan , Lise Getoor , Cristopher Moore

Topic models are widely used in studying social phenomena. We conduct a comparative study examining state-of-the-art neural versus non-neural topic models, performing a rigorous quantitative and qualitative assessment on a dataset of tweets…

计算与语言 · 计算机科学 2021-05-24 Andrew Bennett , Dipendra Misra , Nga Than

Aspiring to achieve an accurate Bitcoin price prediction based on people's opinions on Twitter usually requires millions of tweets, using different text mining techniques (preprocessing, tokenization, stemming, stop word removal), and…

社会与信息网络 · 计算机科学 2023-06-12 Sattarov Otabek , Jaeyoung Choi

To unfold the tremendous amount of multimedia data uploaded daily to social media platforms, effective topic modeling techniques are needed. Existing work tends to apply topic models on written text datasets. In this paper, we propose a…

计算与语言 · 计算机科学 2021-10-29 Lukas Stappen , Jason Thies , Gerhard Hagerer , Björn W. Schuller , Georg Groh

Events of various kinds are mentioned and discussed in text documents, whether they are books, news articles, blogs or microblog feeds. The paper starts by giving an overview of how events are treated in linguistics and philosophy. We…

计算与语言 · 计算机科学 2016-01-18 Jugal Kalita

The amount of user generated contents from various social medias allows analyst to handle a wide view of conversations on several topics related to their business. Nevertheless keeping up-to-date with this amount of information is not…

计算与语言 · 计算机科学 2020-01-31 Jean Valère Cossu , Juan-Manuel Torres-Moreno , Eric SanJuan , Marc El-Bèze

Topic modeling is a useful tool for analyzing large corpora of written documents, particularly academic papers. Despite a wide variety of proposed topic modeling techniques, these techniques do not perform well when applied to medical…

机器学习 · 计算机科学 2025-10-16 Martin Licht , Sara Ketabi , Farzad Khalvati

We use structural topic modeling to examine racial bias in data collected to train models to detect hate speech and abusive language in social media posts. We augment the abusive language dataset by adding an additional feature indicating…

计算与语言 · 计算机科学 2020-05-28 Thomas Davidson , Debasmita Bhattacharya

It has been reported that clustering-based topic models, which cluster high-quality sentence embeddings with an appropriate word selection method, can generate better topics than generative probabilistic topic models. However, these…

计算与语言 · 计算机科学 2023-06-07 Leihang Zhang , Jiapeng Liu , Qiang Yan

Notwithstanding recent work which has demonstrated the potential of using Twitter messages for content-specific data mining and analysis, the depth of such analysis is inherently limited by the scarcity of data imposed by the 140 character…

社会与信息网络 · 计算机科学 2016-11-17 Adham Beykikhoshk , Ognjen Arandjelovic , Dinh Phung , Svetha Venkatesh