中文
相关论文

相关论文: Subject Specific Stream Classification Preprocessi…

200 篇论文

The digital town hall of Twitter becomes a preferred medium of communication for individuals and organizations across the globe. Some of them reach audiences of millions, while others struggle to get noticed. Given the impact of social…

社会与信息网络 · 计算机科学 2021-02-23 Damian Konrad Kowalczyk , Jan Larsen

We explore the feasibility of automatically finding accounts that publish sensitive content on Twitter. One natural approach to this problem is to first create a list of sensitive keywords, and then identify Twitter accounts that use these…

社会与信息网络 · 计算机科学 2017-02-02 Sai Teja Peddinti , Keith W. Ross , Justin Cappos

Data stream algorithms tackle operations on high-volume sequences of read-once data items. Data stream scenarios include inherently real-time systems like sensor networks and financial markets. They also arise in purely-computational…

数据结构与算法 · 计算机科学 2024-03-04 Matthew Andres Moreno , Santiago Rodriguez Papa , Emily Dolson

An identity denotes the role an individual or a group plays in highly differentiated contemporary societies. In this paper, our goal is to classify Twitter users based on their role identities. We first collect a coarse-grained public…

社会与信息网络 · 计算机科学 2020-03-05 Binxuan Huang , Kathleen M. Carley

Cluster analysis is a field of data analysis that extracts underlying patterns in data. One application of cluster analysis is in text-mining, the analysis of large collections of text to find similarities between documents. We used a…

机器学习 · 统计学 2014-08-26 Daniel Godfrey , Caley Johns , Carl Meyer , Shaina Race , Carol Sadek

Research shows that exposure to suicide-related news media content is associated with suicide rates, with some content characteristics likely having harmful and others potentially protective effects. Although good evidence exists for a few…

计算与语言 · 计算机科学 2022-06-29 Hannah Metzler , Hubert Baginski , Thomas Niederkrotenthaler , David Garcia

The manuscript introduces a method to select a random sample from a stream by deciding on each sampling unit immediately after observing it. The process could be applied to unequal as well as equal probability sampling. The implementation…

数据结构与算法 · 计算机科学 2021-11-19 Bardia Panahbehagh , Raphaël Jauslin , Yves Tillé

Data extracted from social networks like Twitter are increasingly being used to build applications and services that mine and summarize public reactions to events, such as traffic monitoring platforms, identification of epidemic outbreaks,…

社会与信息网络 · 计算机科学 2014-05-21 Carlos A. Freitas , Fabrício Benevenuto , Saptarshi Ghosh , Adriano Veloso

The communication revolution has perpetually reshaped the means through which people send and receive information. Social media is an important pillar of this revolution and has brought profound changes to various aspects of our lives.…

In this paper, we introduce the new problem of extracting fine-grained traffic information from Twitter streams by also making publicly available the two (constructed) traffic-related datasets from Belgium and the Brussels capital region.…

计算与语言 · 计算机科学 2021-09-14 Xiangyu Yang , Giannis Bekoulis , Nikos Deligiannis

Twitter has been heavily used as an important channel for communicating and discussing about events in real-time. In such major events, many uninformative tweets are also published rapidly by many users, making it hard to follow the events.…

计算与语言 · 计算机科学 2021-01-15 Renato Stoffalette João

The demand for stream processing is increasing at an unprecedented rate. Big data is no longer limited to processing of big volumes of data. In most real-world scenarios, the need for processing stream data as it comes can only meet the…

分布式、并行与集群计算 · 计算机科学 2018-06-27 Ayush Singhal , Rakesh Pant , Pradeep Sinha

Automatic sentiment analysis play vital role in decision making. Many organizations spend a lot of budget to understand their customer satisfaction by manually going over their feedback/comments or tweets. Automatic sentiment analysis can…

计算与语言 · 计算机科学 2021-07-07 Mohammad Aimal , Maheen Bakhtyar , Junaid Baber , Sadia Lakho , Umar Mohammad , Warda Ahmed , Jahanvash Karim

Mining frequent itemsets through static Databases has been extensively studied and used and is always considered a highly challenging task. For this reason it is interesting to extend it to data streams field. In the streaming case, the…

数据库 · 计算机科学 2012-06-06 Manel Zarrouk , Med Salah Gouider

The role of social media, in particular microblogging platforms such as Twitter, as a conduit for actionable and tactical information during disasters is increasingly acknowledged. However, time-critical analysis of big crisis data on…

计算与语言 · 计算机科学 2016-08-16 Dat Tien Nguyen , Kamela Ali Al Mannai , Shafiq Joty , Hassan Sajjad , Muhammad Imran , Prasenjit Mitra

This paper covers the two approaches for sentiment analysis: i) lexicon based method; ii) machine learning method. We describe several techniques to implement these approaches and discuss how they can be adopted for sentiment classification…

计算与语言 · 计算机科学 2019-02-19 Olga Kolchyna , Tharsis T. P. Souza , Philip Treleaven , Tomaso Aste

Social Networking Sites (SNS) are one of the most important ways of communication. In particular, microblogging sites are being used as analysis avenues due to their peculiarities (promptness, short texts...). There are countless researches…

社会与信息网络 · 计算机科学 2022-06-28 Manuel Francisco , Juan Luis Castro

With the growth of social medias, such as Twitter, plenty of user-generated data emerge daily. The short texts published on Twitter -- the tweets -- have earned significant attention as a rich source of information to guide many…

人工智能 · 计算机科学 2021-06-01 Sérgio Barreto , Ricardo Moura , Jonnathan Carvalho , Aline Paes , Alexandre Plastino

In this paper we present a method to identify tweets that a user may find interesting enough to retweet. The method is based on a global, but personalized classifier, which is trained on data from several users, represented in terms of…

社会与信息网络 · 计算机科学 2017-09-20 Michail Vougioukas , Ion Androutsopoulos , Georgios Paliouras

The literature on data sanitization aims to design algorithms that take an input dataset and produce a privacy-preserving version of it, that captures some of its statistical properties. In this note we study this question from a streaming…

数据结构与算法 · 计算机科学 2021-11-30 Haim Kaplan , Uri Stemmer