English
Related papers

Related papers: Crowdsourcing Dialect Characterization through Twi…

200 papers

We propose an end-to-end neural network to predict the geolocation of a tweet. The network takes as input a number of raw Twitter metadata such as the tweet message and associated user account information. Our model is language independent,…

Computation and Language · Computer Science 2017-10-16 Jey Han Lau , Lianhua Chi , Khoi-Nguyen Tran , Trevor Cohn

Scientific studies investigating laws and regularities of human behavior are nowadays increasingly relying on the wealth of widely available digital information produced by human social activity. In this paper we leverage big data created…

Social and Information Networks · Computer Science 2016-01-22 Stanislav Sobolevsky , Iva Bojic , Alexander Belyi , Izabela Sitko , Bartosz Hawelka , Juan Murillo Arias , Carlo Ratti

Twitter serves as a data source for many Natural Language Processing (NLP) tasks. It can be challenging to identify topics on Twitter due to continuous updating data stream. In this paper, we present an unsupervised graph based framework to…

Computation and Language · Computer Science 2021-04-19 Xiaonan Jing , Qingyuan Hu , Yi Zhang , Julia Taylor Rayz

The present paper uses Twitter to analyze the current state of the worldwide, Spanish-language, independent publishing market. The main purposes are to determine whether certain Latin American Spanish-language independent publishers…

Social and Information Networks · Computer Science 2020-08-04 Ana Gallego-Cuiñas , Esteban Romero-Frías , Wenceslao Arroyo-Machado

Pervasive infrastructures, such as cell phone networks, enable to capture large amounts of human behavioral data but also provide information about the structure of cities and their dynamical properties. In this article, we focus on these…

Citizens are actively interacting with their surroundings, especially through social media. Not only do shared posts give important information about what is happening (from the users' perspective), but also the metadata linked to these…

Social and Information Networks · Computer Science 2023-12-19 Héctor Cerezo-Costas , Ana Fernández Vilas , Manuela Martín-Vicente , Rebeca P. Díaz-Redondo

This paper examines the link between conversational communities on Twitter and their members' expressions of social identity. It specifically tests the presence of community prototypes, or collections of attributes which define a group…

Social and Information Networks · Computer Science 2023-02-21 Thomas Magelinski , Kathleen M. Carley

Recent wide-spread adoption of electronic and pervasive technologies has enabled the study of human behavior at an unprecedented level, uncovering universal patterns underlying human activity, mobility, and inter-personal communication. In…

Physics and Society · Physics 2018-11-21 Alejandro Llorente , Manuel Garcia-Herranz , Manuel Cebrian , Esteban Moro

In this paper, we investigate whether text from a Community Question Answering (QA) platform can be used to predict and describe real-world attributes. We experiment with predicting a wide range of 62 demographic attributes for…

Computation and Language · Computer Science 2017-01-18 Marzieh Saeidi , Alessandro Venerandi , Licia Capra , Sebastian Riedel

Detecting offensive language on Twitter has many applications ranging from detecting/predicting bullying to measuring polarization. In this paper, we focus on building a large Arabic offensive tweet dataset. We introduce a method for…

Computation and Language · Computer Science 2021-03-11 Hamdy Mubarak , Ammar Rashed , Kareem Darwish , Younes Samih , Ahmed Abdelali

Researchers since at least Darwin have debated whether and to what extent emotions are universal or culture-dependent. However, previous studies have primarily focused on facial expressions and on a limited set of emotions. Given that…

Computation and Language · Computer Science 2013-04-30 Eugene Yuta Bann , Joanna J. Bryson

We present a new computational technique to detect and analyze statistically significant geographic variation in language. Our meta-analysis approach captures statistical properties of word usage across geographical regions and uses…

Computation and Language · Computer Science 2016-03-08 Vivek Kulkarni , Bryan Perozzi , Steven Skiena

To analyse large numbers of texts, social science researchers are increasingly confronting the challenge of text classification. When manual labeling is not possible and researchers have to find automatized ways to classify texts, computer…

Computation and Language · Computer Science 2023-10-10 Karina Shyrokykh , Maksym Girnyk , Lisa Dellmuth

Identifying the language of social media messages is an important first step in linguistic processing. Existing models for Twitter focus on content analysis, which is successful for dissimilar language pairs. We propose a label propagation…

Computation and Language · Computer Science 2016-07-20 Will Radford , Matthias Galle

This paper addresses the task of user gender classification in social media, with an application to Twitter. The approach automatically predicts gender by leveraging observable information such as the tweet behavior, linguistic content of…

Information Retrieval · Computer Science 2014-05-27 Puneet Singh Ludu

Topic modeling is a key method in text analysis, but existing approaches fail to efficiently scale to large datasets or are limited by assuming one topic per document. Overcoming these limitations, we introduce Semantic Component Analysis…

Computation and Language · Computer Science 2025-09-29 Florian Eichin , Carolin M. Schuster , Georg Groh , Michael A. Hedderich

Variations in languages across geographic regions or cultures are crucial to address to avoid biases in NLP systems designed for culturally sensitive tasks, such as hate speech detection or dialog with conversational agents. In languages…

Computation and Language · Computer Science 2024-12-17 Javier A. Lopetegui , Arij Riabi , Djamé Seddah

Mixed language data is one of the difficult yet less explored domains of natural language processing. Most research in fields like machine translation or sentiment analysis assume monolingual input. However, people who are capable of using…

Neural and Evolutionary Computing · Computer Science 2014-12-23 Joseph Chee Chang , Chu-Cheng Lin

This paper presents our approach for SwissText & KONVENS 2020 shared task 2, which is a multi-stage neural model for Swiss German (GSW) identification on Twitter. Our model outputs either GSW or non-GSW and is not meant to be used as a…

Computation and Language · Computer Science 2020-06-08 Mohammadreza Banaei , Rémi Lebret , Karl Aberer

This work investigates style and topic aspects of language in online communities: looking at both utility as an identifier of the community and correlation with community reception of content. Style is characterized using a hybrid word and…

Computation and Language · Computer Science 2016-09-16 Trang Tran , Mari Ostendorf