中文
相关论文

相关论文: Impact of Feature Selection on Micro-Text Classifi…

200 篇论文

Sentiment analysis of microblogs such as Twitter has recently gained a fair amount of attention. One of the simplest sentiment analysis approaches compares the words of a posting against a labeled word list, where each word has been scored…

信息检索 · 计算机科学 2017-10-12 Finn Årup Nielsen

Using machine learning algorithms, including deep learning, we studied the prediction of personal attributes from the text of tweets, such as gender, occupation, and age groups. We applied word2vec to construct word vectors, which were then…

计算机与社会 · 计算机科学 2017-12-27 Take Yo , Kazutoshi Sasahara

In recent years, social bots have been using increasingly more sophisticated, challenging detection strategies. While many approaches and features have been proposed, social bots evade detection and interact much like humans making it…

社会与信息网络 · 计算机科学 2018-12-20 Isa Inuwa-Dutse , Bello Shehu Bello , Ioannis Korkontzelos

Twitter messages often contain so-called hashtags to denote keywords related to them. Using a dataset of 29 million messages, I explore relations among these hashtags with respect to co-occurrences. Furthermore, I present an attempt to…

计算与语言 · 计算机科学 2011-11-29 Jan Pöschko

Research shows that exposure to suicide-related news media content is associated with suicide rates, with some content characteristics likely having harmful and others potentially protective effects. Although good evidence exists for a few…

计算与语言 · 计算机科学 2022-06-29 Hannah Metzler , Hubert Baginski , Thomas Niederkrotenthaler , David Garcia

Twitter is one of the most popular microblogging services in the world. The great amount of information within Twitter makes it an important information channel for people to learn and share news. Twitter hashtag is an popular feature that…

社会与信息网络 · 计算机科学 2018-05-01 Shih-Feng Yang , Julia Taylor Rayz

The role of social media, in particular microblogging platforms such as Twitter, as a conduit for actionable and tactical information during disasters is increasingly acknowledged. However, time-critical analysis of big crisis data on…

计算与语言 · 计算机科学 2016-08-16 Dat Tien Nguyen , Kamela Ali Al Mannai , Shafiq Joty , Hassan Sajjad , Muhammad Imran , Prasenjit Mitra

Applying natural language processing for mining and intelligent information access to tweets (a form of microblog) is a challenging, emerging research area. Unlike carefully authored news text and other longer content, tweets pose a number…

The sequence of documents produced by any given author varies in style and content, but some documents are more typical or representative of the source than others. We quantify the extent to which a given short text is characteristic of a…

计算与语言 · 计算机科学 2019-09-10 Charuta Pethe , Steven Skiena

Twitter data has been shown broadly applicable for public health surveillance. Previous public health studies based on Twitter data have largely relied on keyword-matching or topic models for clustering relevant tweets. However, both…

计算与语言 · 计算机科学 2019-12-04 Xiaoyi Zhang , Rodoniki Athanasiadou , Narges Razavian

Text from social media provides a set of challenges that can cause traditional NLP approaches to fail. Informal language, spelling errors, abbreviations, and special characters are all commonplace in these posts, leading to a prohibitively…

机器学习 · 计算机科学 2016-05-18 Bhuwan Dhingra , Zhong Zhou , Dylan Fitzpatrick , Michael Muehl , William W. Cohen

Word embeddings and convolutional neural networks (CNN) have attracted extensive attention in various classification tasks for Twitter, e.g. sentiment classification. However, the effect of the configuration used to train and generate the…

信息检索 · 计算机科学 2017-03-23 Xiao Yang , Craig Macdonald , Iadh Ounis

How can we study social interactions on evolving topics at a mass scale? Over the past decade, researchers from diverse fields such as economics, political science, and public health have often done this by querying Twitter's public API…

社会与信息网络 · 计算机科学 2022-09-23 Sacha Lévy , Farimah Poursafaei , Kellin Pelrine , Reihaneh Rabbany

Extracting topics from large collections of unstructured text-documents has become a central task in current NLP applications and algorithms like NMF, LDA as well as their generalizations are the well-established current state of the art.…

社会与信息网络 · 计算机科学 2021-11-23 Mattias Luber , Anton Thielmann , Christoph Weisser , Benjamin Säfken

The micro-blogging platform Twitter allows its nearly 320 million monthly active users to build a network of follower connections to other Twitter users (i.e., followees) in order to subscribe to content posted by these users. With this…

信息检索 · 计算机科学 2018-09-11 Dominik Kowald , Elisabeth Lex

Cluster analysis is a field of data analysis that extracts underlying patterns in data. One application of cluster analysis is in text-mining, the analysis of large collections of text to find similarities between documents. We used a…

机器学习 · 统计学 2014-08-26 Daniel Godfrey , Caley Johns , Carl Meyer , Shaina Race , Carol Sadek

In this paper, we explore meta-learning for few-shot text classification. Meta-learning has shown strong performance in computer vision, where low-level patterns are transferable across learning tasks. However, directly applying this…

计算与语言 · 计算机科学 2020-02-19 Yujia Bao , Menghua Wu , Shiyu Chang , Regina Barzilay

In microblogging, hashtags are used to be topical markers, and they are adopted by users that contribute similar content or express a related idea. However, hashtags are created in a free style and there is no domain category information…

信息检索 · 计算机科学 2015-05-06 Shuangyong Song , Yao Meng

Detecting harmful content on social media, such as Twitter, is made difficult by the fact that the seemingly simple yes/no classification conceals a significant amount of complexity. Unfortunately, while several datasets have been collected…

计算与语言 · 计算机科学 2023-11-14 Saad Almohaimeed , Saleh Almohaimeed , Ashfaq Ali Shafin , Bogdan Carbunar , Ladislau Bölöni

Semantic sentence embeddings are usually supervisedly built minimizing distances between pairs of embeddings of sentences labelled as semantically similar by annotators. Since big labelled datasets are rare, in particular for non-English…

计算与语言 · 计算机科学 2021-10-06 Marco Di Giovanni , Marco Brambilla