English
Related papers

Related papers: Learning about Spanish dialects through Twitter

200 papers

Our entry into the HAHA 2019 Challenge placed $3^{rd}$ in the classification task and $2^{nd}$ in the regression task. We describe our system and innovations, as well as comparing our results to a Naive Bayes baseline. A large Twitter based…

Computation and Language · Computer Science 2019-07-09 Bobak Farzin , Piotr Czapla , Jeremy Howard

Much work in the space of NLP has used computational methods to explore sociolinguistic variation in text. In this paper, we argue that memes, as multimodal forms of language comprised of visual templates and text, also exhibit meaningful…

Computation and Language · Computer Science 2023-11-16 Naitian Zhou , David Jurgens , David Bamman

Social media is considered a democratic space in which people connect and interact with each other regardless of their gender, race, or any other demographic aspect. Despite numerous efforts that explore demographic aspects in social media,…

Social and Information Networks · Computer Science 2018-04-03 Johnnatan Messias

We propose an end-to-end neural network to predict the geolocation of a tweet. The network takes as input a number of raw Twitter metadata such as the tweet message and associated user account information. Our model is language independent,…

Computation and Language · Computer Science 2017-10-16 Jey Han Lau , Lianhua Chi , Khoi-Nguyen Tran , Trevor Cohn

Political polarization generates strong effects on society, driving controversial debates and influencing the institutions. Territorial disputes are one of the most important polarized scenarios and have been consistently related to the use…

Physics and Society · Physics 2026-02-25 Julia Atienza-Barthelemy , Samuel Martin-Gutierrez , Juan C. Losada , Rosa M. Benito

This paper presents a computational approach to author profiling taking gender and language variety into account. We apply an ensemble system with the output of multiple linear SVM classifiers trained on character and word $n$-grams. We…

Computation and Language · Computer Science 2017-07-04 Alina Maria Ciobanu , Marcos Zampieri , Shervin Malmasi , Liviu P. Dinu

Languages can encode temporal subordination lexically, via subordinating conjunctions, and morphologically, by marking the relation on the predicate. Systematic cross-linguistic variation among the former can be studied using…

Computation and Language · Computer Science 2024-07-31 Nilo Pedrazzini

Our paper studies the predictability of online speech -- that is, how well language models learn to model the distribution of user generated content on X (previously Twitter). We define predictability as a measure of the model's…

Computation and Language · Computer Science 2026-01-07 Mina Remeli , Moritz Hardt , Robert C. Williamson

In recent years, multimodal natural language processing, aimed at learning from diverse data types, has garnered significant attention. However, there needs to be more clarity when it comes to analysing multimodal tasks in multi-lingual…

Computation and Language · Computer Science 2024-06-13 Gaurish Thakkar , Sherzod Hakimov , Marko Tadić

Language change is influenced by many factors, but often starts from synchronic variation, where multiple linguistic patterns or forms coexist, or where different speech communities use language in increasingly different ways. Besides…

Social and Information Networks · Computer Science 2023-09-06 Andres Karjus , Christine Cuskley

Profiting from the emergence of web-scale social data sets, numerous recent studies have systematically explored human mobility patterns over large populations and large time scales. Relatively little attention, however, has been paid to…

On-line social networks have grown quickly over the last few years and nowadays many people use them frequently. Furthermore the emergence of smartphones allows to access these networks any time from any physical location. Among the social…

Social and Information Networks · Computer Science 2014-04-29 Antònia Tugores , Pere Colet

Despite its importance, the time variable has been largely neglected in the NLP and language model literature. In this paper, we present TimeLMs, a set of language models specialized on diachronic Twitter data. We show that a continual…

Computation and Language · Computer Science 2022-04-04 Daniel Loureiro , Francesco Barbieri , Leonardo Neves , Luis Espinosa Anke , Jose Camacho-Collados

We observe and report on a systematic relationship between population density and Twitter use. Number of tweets, number of users and population per unit area are related by power laws, with exponents greater than one, that are consistent…

Social and Information Networks · Computer Science 2017-11-28 Rudy Arthur , Hywel Williams

This paper demonstrates a two-stage method for deriving insights from social media data relating to disinformation by applying a combination of geospatial classification and embedding-based language modelling across multiple languages. In…

Computation and Language · Computer Science 2021-08-09 David Tuxworth , Dimosthenis Antypas , Luis Espinosa-Anke , Jose Camacho-Collados , Alun Preece , David Rogers

Several computational models have been developed to detect and analyze dialect variation in recent years. Most of these models assume a predefined set of geographical regions over which they detect and analyze dialectal variation. However,…

Computation and Language · Computer Science 2019-10-17 Hang Jiang , Haoshen Hong , Yuxing Chen , Vivek Kulkarni

Complex networks are important tools for analyzing the information flow in many aspects of nature and human society. Using data from the microblogging service Twitter, we study networks of correlations in the appearance of words from three…

Physics and Society · Physics 2013-10-23 Joachim Mathiesen , Pernilly Yde , Mogens H. Jensen

Geolocalization of social media content is the task of determining the geographical location of a user based on textual data, that may show linguistic variations and informal language. In this project, we address the GeoLingIt challenge of…

Computation and Language · Computer Science 2024-07-24 Davide Savarro , Davide Zago , Stefano Zoia

Many aspects of people's lives are proven to be deeply connected to their jobs. In this paper, we first investigate the distinct characteristics of major occupation categories based on tweets. From multiple social media platforms, we gather…

Computers and Society · Computer Science 2017-01-24 Tianran Hu , Haoyuan Xiao , Thuy-vy Thi Nguyen , Jiebo Luo

This work presents a set of experiments conducted to predict the gender of Twitter users based on language-independent features extracted from the text of the users' tweets. The experiments were performed on a version of TwiSty dataset…

Computation and Language · Computer Science 2024-12-02 Reyhaneh Hashempour , Barbara Plank , Aline Villavicencio , Renato Cordeiro de Amorim
‹ Prev 1 3 4 5 6 7 10 Next ›