中文
相关论文

相关论文: Distant Supervision for Topic Classification of Tw…

200 篇论文

Twitter serves as a data source for many Natural Language Processing (NLP) tasks. It can be challenging to identify topics on Twitter due to continuous updating data stream. In this paper, we present an unsupervised graph based framework to…

计算与语言 · 计算机科学 2021-04-19 Xiaonan Jing , Qingyuan Hu , Yi Zhang , Julia Taylor Rayz

Social media channels such as Twitter have emerged as popular platforms for crowds to respond to public events such as speeches, sports and debates. While this promises tremendous opportunities to understand and make sense of the reception…

机器学习 · 计算机科学 2012-12-24 Yuheng Hu , Ajita John , Fei Wang , Doree Duncan Seligmann , Subbarao Kambhampati

Millions of people express themselves on public social media, such as Twitter. Through their posts, these people may reveal themselves as potentially valuable sources of information. For example, real-time information about an event might…

社会与信息网络 · 计算机科学 2014-04-09 Jalal Mahmud , Michelle Zhou , Nimrod Megiddo , Jeffrey Nichols , Clemens Drews

Discussions on Twitter involve participation from different communities with different dialects and it is often necessary to summarize a large number of posts into a representative sample to provide a synopsis. Yet, any such representative…

计算机与社会 · 计算机科学 2021-04-06 Vijay Keswani , L. Elisa Celis

In the past few years, there has been a huge growth in Twitter sentiment analysis having already provided a fair amount of research on sentiment detection of public opinion among Twitter users. Given the fact that Twitter messages are…

Polarized topics often spark discussion and debate on social media. Recent studies have shown that polarized debates have a specific clustered structure in the endorsement net- work, which indicates that users direct their endorsements…

社会与信息网络 · 计算机科学 2017-04-03 Kiran Garimella , Gianmarco De Francisci Morales , Aristides Gionis , Michael Mathioudakis

As a YouTube channel grows, each video can potentially collect enormous amounts of comments that provide direct feedback from the viewers. These comments are a major means of understanding viewer expectations and improving channel…

信息检索 · 计算机科学 2023-06-05 Rhitabrat Pokharel , Dixit Bhatta

We propose a streaming algorithm for the binary classification of data based on crowdsourcing. The algorithm learns the competence of each labeller by comparing her labels to those of other labellers on the same tasks and uses this…

机器学习 · 统计学 2016-02-24 Thomas Bonald , Richard Combes

Topic modeling is a key component in unsupervised learning, employed to identify topics within a corpus of textual data. The rapid growth of social media generates an ever-growing volume of textual data daily, making online topic modeling…

机器学习 · 计算机科学 2025-10-23 Federica Granese , Benjamin Navet , Serena Villata , Charles Bouveyron

We aim at solving the problem of predicting people's ideology, or political tendency. We estimate it by using Twitter data, and formalize it as a classification problem. Ideology-detection has long been a challenging yet important problem.…

机器学习 · 计算机科学 2020-06-19 Zhiping Xiao , Weiping Song , Haoyan Xu , Zhicheng Ren , Yizhou Sun

The real-time nature of Twitter means that term distributions in tweets and in search queries change rapidly: the most frequent terms in one hour may look very different from those in the next. Informally, we call this phenomenon "churn".…

信息检索 · 计算机科学 2012-06-01 Jimmy Lin , Gilad Mishne

Information quality in social media is an increasingly important issue, but web-scale data hinders experts' ability to assess and correct much of the inaccurate content, or `fake news,' present in these platforms. This paper develops a…

社会与信息网络 · 计算机科学 2018-06-01 Cody Buntain , Jennifer Golbeck

Discourse parsing, the task of analyzing the internal rhetorical structure of texts, is a challenging problem in natural language processing. Despite the recent advances in neural models, the lack of large-scale, high-quality corpora for…

计算与语言 · 计算机科学 2023-05-24 Feng Jiang , Longwang He , Peifeng Li , Qiaoming Zhu , Haizhou Li

Most previous work related to tweet classification have focused on identifying a given tweet as a spam, or to classify a Twitter user account as a spammer or a bot. In most cases the tweet classification has taken place offline, on a…

社会与信息网络 · 计算机科学 2018-02-06 Jonas Lundberg , Jonas Nordqvist , Antonio Matosevic

Twitter introduced user lists in late 2009, allowing users to be grouped according to meaningful topics or themes. Lists have since been adopted by media outlets as a means of organising content around news stories. Thus the curation of…

社会与信息网络 · 计算机科学 2012-07-03 Derek Greene , Gavin Sheridan , Barry Smyth , Pádraig Cunningham

Various domain users are increasingly leveraging real-time social media data to gain rapid situational awareness. However, due to the high noise in the deluge of data, effectively determining semantically relevant information can be…

社会与信息网络 · 计算机科学 2019-10-09 Luke S. Snyder , Yi-Shan Lin , Morteza Karimzadeh , Dan Goldwasser , David S. Ebert

Rumors are rampant in the era of social media. Conversation structures provide valuable clues to differentiate between real and fake claims. However, existing rumor detection methods are either limited to the strict relation of user…

计算与语言 · 计算机科学 2021-11-16 Hongzhan Lin , Jing Ma , Mingfei Cheng , Zhiwei Yang , Liangliang Chen , Guang Chen

For more than a decade now, academicians and online platform administrators have been studying solutions to the problem of bot detection. Bots are computer algorithms whose use is far from being benign: malicious bots are purposely created…

密码学与安全 · 计算机科学 2025-06-25 Rocco De Nicola , Marinella Petrocchi , Manuel Pratelli

Distant supervision is a popular method for performing relation extraction from text that is known to produce noisy labels. Most progress in relation extraction and classification has been made with crowdsourced corrections to…

计算与语言 · 计算机科学 2022-09-21 Anca Dumitrache , Lora Aroyo , Chris Welty

We propose a novel training and inference method for detecting political bias in long text content such as newspaper opinion articles. Obtaining long text data and annotations at sufficient scale for training is difficult, but it is…

计算与语言 · 计算机科学 2019-11-20 Aditya Saligrama