English
Related papers

Related papers: Discriminating between similar languages in Twitte…

200 papers

Social media is considered a democratic space in which people connect and interact with each other regardless of their gender, race, or any other demographic aspect. Despite numerous efforts that explore demographic aspects in social media,…

Social and Information Networks · Computer Science 2018-04-03 Johnnatan Messias

Fake news detection in social media has become increasingly important due to the rapid proliferation of personal media channels and the consequential dissemination of misleading information. Existing methods, which primarily rely on…

Multimedia · Computer Science 2024-06-17 Wanqing Zhao , Yuta Nakashima , Haiyuan Chen , Noboru Babaguchi

In this paper we show how the performance of tweet clustering can be improved by leveraging character-based neural networks. The proposed approach overcomes the limitations related to the vocabulary explosion in the word-based models and…

Information Retrieval · Computer Science 2017-03-17 Svitlana Vakulenko , Lyndon Nixon , Mihai Lupu

We propose a new computational approach for tracking and detecting statistically significant linguistic shifts in the meaning and usage of words. Such linguistic shifts are especially prevalent on the Internet, where the rapid exchange of…

Computation and Language · Computer Science 2014-11-13 Vivek Kulkarni , Rami Al-Rfou , Bryan Perozzi , Steven Skiena

Our paper studies the predictability of online speech -- that is, how well language models learn to model the distribution of user generated content on X (previously Twitter). We define predictability as a measure of the model's…

Computation and Language · Computer Science 2026-01-07 Mina Remeli , Moritz Hardt , Robert C. Williamson

An identity denotes the role an individual or a group plays in highly differentiated contemporary societies. In this paper, our goal is to classify Twitter users based on their role identities. We first collect a coarse-grained public…

Social and Information Networks · Computer Science 2020-03-05 Binxuan Huang , Kathleen M. Carley

This work presents a supervised method for generating a classifier model of the stances held by Chinese-speaking politicians and other Twitter users. Many previous works of political tweets prediction exist on English tweets, but to the…

Computers and Society · Computer Science 2021-10-13 Fenglei Gu , Duoji Jiang

In applications involving conversational speech, data sparsity is a limiting factor in building a better language model. We propose a simple, language-independent method to quickly harvest large amounts of data from Twitter to supplement a…

Computation and Language · Computer Science 2015-04-13 Aaron Jaech , Mari Ostendorf

In this paper we examine methods to detect hate speech in social media, while distinguishing this from general profanity. We aim to establish lexical baselines for this task by applying supervised classification methods using a recently…

Computation and Language · Computer Science 2017-12-29 Shervin Malmasi , Marcos Zampieri

Providing better language tools for low-resource and endangered languages is imperative for equitable growth. Recent progress with massively multilingual pretrained models has proven surprisingly effective at performing zero-shot transfer…

Computation and Language · Computer Science 2022-11-10 Louis Clouâtre , Prasanna Parthasarathi , Amal Zouaq , Sarath Chandar

Semantic sentence embeddings are usually supervisedly built minimizing distances between pairs of embeddings of sentences labelled as semantically similar by annotators. Since big labelled datasets are rare, in particular for non-English…

Computation and Language · Computer Science 2021-10-06 Marco Di Giovanni , Marco Brambilla

Variation in language is ubiquitous, particularly in newer forms of writing such as social media. Fortunately, variation is not random, it is often linked to social properties of the author. In this paper, we show how to exploit social…

Computation and Language · Computer Science 2017-08-29 Yi Yang , Jacob Eisenstein

In this paper we present results from a method of community detection using label propagation in undirected, unweighted graphs which incorporates elements of neural computing and spike-based data. Using a fully connected, edge-weighted…

Disordered Systems and Neural Networks · Physics 2018-10-24 Kathleen E. Hamilton , Travis S. Humble

The rise in popularity and ubiquity of Twitter has made sentiment analysis of tweets an important and well-covered area of research. However, the 140 character limit imposed on tweets makes it hard to use standard linguistic methods for…

Social and Information Networks · Computer Science 2021-01-05 Soroush Vosoughi , Helen Zhou , Deb Roy

Twitter continuously tightens the access to its data via the publicly accessible, cost-free standard APIs. This especially applies to the follow network. In light of this, we successfully modified a network sampling method to work…

Social and Information Networks · Computer Science 2021-01-13 Felix Victor Münch , Ben Thies , Cornelius Puschmann , Axel Bruns

With the rise in popularity of public social media and micro-blogging services, most notably Twitter, the people have found a venue to hear and be heard by their peers without an intermediary. As a consequence, and aided by the public…

Computation and Language · Computer Science 2016-06-21 Prashanth Vijayaraghavan , Soroush Vosoughi , Deb Roy

Text actionability detection is the problem of classifying user authored natural language text, according to whether it can be acted upon by a responding agent. In this paper, we propose a supervised learning framework for domain-aware,…

Information Retrieval · Computer Science 2015-11-04 Nemanja Spasojevic , Adithya Rao

Identifying clusters or community structures in networks has become an integral part of social network analysis. Though many methods were proposed, the label propagation algorithm (LPA) is a popular computationally efficient method with…

Social and Information Networks · Computer Science 2022-01-19 Jyothimon chandran , Madhuviswanatham Vankadara

On daily basis, millions of Twitter accounts post a vast number of tweets including numerous Twitter entities (mentions, replies, hashtags, photos, URLs). Many of these entities are used in common by many accounts. The more common entities…

Social and Information Networks · Computer Science 2015-06-02 Gerasimos Razis , Ioannis Anagnostopoulos

Part-of-Speech (POS) tagging is an important component of the NLP pipeline, but many low-resource languages lack labeled data for training. An established method for training a POS tagger in such a scenario is to create a labeled training…

Computation and Language · Computer Science 2022-11-01 Ayyoob Imani , Silvia Severini , Masoud Jalili Sabet , François Yvon , Hinrich Schütze