中文
相关论文

相关论文: Microblog Topic Identification using Linked Open D…

200 篇论文

Protection of human rights is one of the most important problems of our world. In this paper, our aim is to provide a dataset which covers one of the most significant human rights contradiction in recent months affected the whole world,…

计算与语言 · 计算机科学 2023-10-18 Hasan Kemik , Nusret Özateş , Meysam Asgari-Chenaghlu , Yang Li , Erik Cambria

As social networks are constantly changing and evolving, methods to analyze dynamic social networks are becoming more important in understanding social trends. However, due to the restrictions imposed by the social network service…

社会与信息网络 · 计算机科学 2018-01-09 Kaan Bingöl , Bahaeddin Eravcı , Çağrı Özgenç Etemoğlu , Hakan Ferhatosmanoğlu , Buğra Gedik

Understanding causal language in informal discourse is a core yet underexplored challenge in NLP. Existing datasets largely focus on explicit causality in structured text, providing limited support for detecting implicit causal expressions,…

计算与语言 · 计算机科学 2025-09-23 Xiaohan Ding , Kaike Ping , Buse Çarık , Eugenia Rho

This study introduces and investigates the capabilities of three different text mining approaches, namely Latent Semantic Analysis, Latent Dirichlet Analysis, and Clustering Word Vectors, for automating code extraction from a relatively…

机器学习 · 计算机科学 2023-04-20 Sina Mahdipour Saravani , Sadaf Ghaffari , Yanye Luther , James Folkestad , Marcia Moraes

In the era of data-driven journalism, data analytics can deliver tools to support journalists in connecting to new and developing news stories, e.g., as echoed in micro-blogs such as Twitter, the new citizen-driven media. In this paper, we…

社会与信息网络 · 计算机科学 2014-05-14 Bichen Shi , Georgiana Ifrim , Neil Hurley

The article describes the approaches for forming different predictive features of tweet data sets and using them in the predictive analysis for decision-making support. The graph theory as well as frequent itemsets and association rules…

计算与语言 · 计算机科学 2022-01-07 Bohdan M. Pavlyshenko

Rapid expansion of social media platforms such as X (formerly Twitter), Facebook, and Reddit has enabled large-scale analysis of public perceptions on diverse topics, including social issues, politics, natural disasters, and consumer…

计算与语言 · 计算机科学 2025-12-09 Aoi Fujita , Taichi Yamamoto , Yuri Nakayama , Ryota Kobayashi

Twitter serves as a data source for many Natural Language Processing (NLP) tasks. It can be challenging to identify topics on Twitter due to continuous updating data stream. In this paper, we present an unsupervised graph based framework to…

计算与语言 · 计算机科学 2021-04-19 Xiaonan Jing , Qingyuan Hu , Yi Zhang , Julia Taylor Rayz

Supervised topic models can help clinical researchers find interpretable cooccurence patterns in count data that are relevant for diagnostics. However, standard formulations of supervised Latent Dirichlet Allocation have two problems.…

Although latent factor models (e.g., matrix factorization) obtain good performance in predictions, they suffer from several problems including cold-start, non-transparency, and suboptimal recommendations. In this paper, we employ text with…

机器学习 · 计算机科学 2022-03-03 Biyi Fang , Kripa Rajshekhar , Diego Klabjan

Hashtags are semantico-syntactic constructs used across various social networking and microblogging platforms to enable users to start a topic specific discussion or classify a post into a desired category. Segmenting and linking the…

信息检索 · 计算机科学 2015-01-15 Piyush Bansal , Romil Bansal , Vasudeva Varma

Topic detection is a challenging task, especially without knowing the exact number of topics. In this paper, we present a novel approach based on neural network to detect topics in the micro-blogging dataset. We use an unsupervised neural…

信息检索 · 计算机科学 2020-06-18 Cong Wan , Shan Jiang , Cuirong Wang , Cong Wang , Changming Xu , Xianxia Chen , Ying Yuan

Contemporary datasets on tobacco consumption focus on one of two topics, either public health mentions and disease surveillance, or sentiment analysis on topical tobacco products and services. However, two primary considerations are not…

计算与语言 · 计算机科学 2020-06-16 Kartikey Pant , Venkata Himakar Yanamandra , Alok Debnath , Radhika Mamidi

The widespread use of offensive content in social media has led to an abundance of research in detecting language such as hate speech, cyberbullying, and cyber-aggression. Recent work presented the OLID dataset, which follows a taxonomy for…

计算与语言 · 计算机科学 2021-09-27 Sara Rosenthal , Pepa Atanasova , Georgi Karadzhov , Marcos Zampieri , Preslav Nakov

Social media platforms such as Twitter (now X) provide rich data for analyzing public discourse, especially during crises such as the COVID-19 pandemic. However, the brevity, informality, and noise of social media short texts often hinder…

计算与语言 · 计算机科学 2025-10-23 Wangjiaxuan Xin , Shuhua Yin , Shi Chen , Yaorong Ge

Sentiment analysis possesses the potential of diverse applicability on digital platforms. Sentiment analysis extracts the polarity to understand the intensity and subjectivity in the text. This work uses a lexicon-based method to perform…

计算与语言 · 计算机科学 2024-09-20 Muhammad Raees , Samina Fazilat

Twitter is among the most prevalent social media platform being used by millions of people all over the world. It is used to express ideas and opinions about political, social, business, sports, health, religion, and various other…

计算与语言 · 计算机科学 2021-12-07 Khubaib Ahmed Qureshi

Politics is one of the most prevalent topics discussed on social media platforms, particularly during major election cycles, where users engage in conversations about candidates and electoral processes. Malicious actors may use this…

计算与语言 · 计算机科学 2024-04-29 Alphaeus Dmonte , Marcos Zampieri , Kevin Lybarger , Massimiliano Albanese , Genya Coulter

Thematic analysis of social media posts provides a major understanding of public discourse, yet traditional methods often struggle to capture the complexity and nuance of unstructured, large-scale text data. This study introduces a novel…

计算与语言 · 计算机科学 2025-03-05 Mohammed-Khalil Ghali , Abdelrahman Farrag , Sarah Lam , Daehan Won

Topic modeling is a key method in text analysis, but existing approaches fail to efficiently scale to large datasets or are limited by assuming one topic per document. Overcoming these limitations, we introduce Semantic Component Analysis…

计算与语言 · 计算机科学 2025-09-29 Florian Eichin , Carolin M. Schuster , Georg Groh , Michael A. Hedderich