English
Related papers

Related papers: Learning about Spanish dialects through Twitter

200 papers

The massive popularity of online social media provides a unique opportunity for researchers to study the linguistic characteristics and patterns of user's interactions. In this paper, we provide an in-depth characterization of language…

Social and Information Networks · Computer Science 2017-05-12 Pantelis Vikatos , Johnnatan Messias , Manoel Miranda , Fabricio Benevenuto

In this paper we take into account both social and linguistic aspects to perform demographic analysis by processing a large amount of tweets in Basque language. The study of demographic characteristics and social relationships are…

Computers and Society · Computer Science 2021-09-09 J. Fernandez de Landa , R. Agerri

This study presents an LLM-assisted annotation pipeline for the sociolinguistic and topical analysis of bilingual discourse in two typologically distinct contexts: Spanish-English and Spanish-Guaran\'i. Using large language models, we…

Computation and Language · Computer Science 2025-12-04 Nemika Tyagi , Nelvin Licona Guevara , Olga Kellert

In recent years, large language models (LLMs) have demonstrated a high capacity for understanding and generating text in Spanish. However, with five hundred million native speakers, Spanish is not a homogeneous language but rather one rich…

Computation and Language · Computer Science 2025-05-22 Marina Mayor-Rocher , Cristina Pozo , Nina Melero , Gonzalo Martínez , María Grandury , Pedro Reviriego

We investigate the predictive power behind the language of food on social media. We collect a corpus of over three million food-related posts from Twitter and demonstrate that many latent population characteristics can be directly predicted…

Computation and Language · Computer Science 2016-11-15 Daniel Fried , Mihai Surdeanu , Stephen Kobourov , Melanie Hingle , Dane Bell

The prevalence of social media presents a growing opportunity to collect and analyse examples of English varieties. Whilst usage of these varieties was - and, in many cases, still is - used only in spoken contexts or hard-to-access private…

Computation and Language · Computer Science 2024-01-23 Nhi Pham , Lachlan Pham , Adam L. Meyers

We propose a new computational approach for tracking and detecting statistically significant linguistic shifts in the meaning and usage of words. Such linguistic shifts are especially prevalent on the Internet, where the rapid exchange of…

Computation and Language · Computer Science 2014-11-13 Vivek Kulkarni , Rami Al-Rfou , Bryan Perozzi , Steven Skiena

As global political preeminence gradually shifted from the United Kingdom to the United States, so did the capacity to culturally influence the rest of the world. In this work, we analyze how the world-wide varieties of written English are…

Computation and Language · Computer Science 2018-05-29 Bruno Gonçalves , Lucía Loureiro-Porto , José J. Ramasco , David Sánchez

We propose a method for embedding two-dimensional locations in a continuous vector space using a neural network-based model incorporating mixtures of Gaussian distributions, presenting two model variants for text-based geolocation and…

Computation and Language · Computer Science 2017-08-16 Afshin Rahimi , Timothy Baldwin , Trevor Cohn

Dictionaries and phrase tables are the basis of modern statistical machine translation systems. This paper develops a method that can automate the process of generating and extending dictionaries and phrase tables. Our method can translate…

Computation and Language · Computer Science 2013-09-18 Tomas Mikolov , Quoc V. Le , Ilya Sutskever

The appearance of complex attention-based language models such as BERT, Roberta or GPT-3 has allowed to address highly complex tasks in a plethora of scenarios. However, when applied to specific domains, these models encounter considerable…

Computation and Language · Computer Science 2022-06-14 Javier Huertas-Tato , Alejandro Martin , David Camacho

Typological databases in linguistics are usually categorical-valued. As a result, it is difficult to have a clear visualization of the data. In this paper, we describe a workflow to analyze the topological shapes of South American languages…

Computation and Language · Computer Science 2024-03-26 Rui Dong

We present a new corpus of Twitter data annotated for codeswitching and borrowing between Spanish and English. The corpus contains 9,500 tweets annotated at the token level with codeswitches, borrowings, and named entities. This corpus…

Computation and Language · Computer Science 2022-06-13 Elena Alvarez Mellado , Constantine Lignos

Identifying the language of social media messages is an important first step in linguistic processing. Existing models for Twitter focus on content analysis, which is successful for dissimilar language pairs. We propose a label propagation…

Computation and Language · Computer Science 2016-07-20 Will Radford , Matthias Galle

This paper introduces how human languages can be studied in light of recent development of network theories. There are two directions of exploration. One is to study networks existing in the language system. Various lexical networks can be…

Computation and Language · Computer Science 2007-05-23 Jinyun KE

Word embeddings provide an unsupervised way to understand differences in word usage between discursive communities. A number of recent papers have focused on identifying words that are used differently by two or more communities. But word…

Computation and Language · Computer Science 2023-02-14 Thyge Enggaard , August Lohse , Morten Axel Pedersen , Sune Lehmann

We investigate the behavior of maps learned by machine translation methods. The maps translate words by projecting between word embedding spaces of different languages. We locally approximate these maps using linear maps, and find that they…

Computation and Language · Computer Science 2018-06-19 Ndapa Nakashole , Raphael Flauger

We propose an LSTM-based model with hierarchical architecture on named entity recognition from code-switching Twitter data. Our model uses bilingual character representation and transfer learning to address out-of-vocabulary words. In order…

Computation and Language · Computer Science 2019-06-11 Genta Indra Winata , Chien-Sheng Wu , Andrea Madotto , Pascale Fung

Sentiment analysis is a text mining task that determines the polarity of a given text, i.e., its positiveness or negativeness. Recently, it has received a lot of attention given the interest in opinion mining in micro-blogging platforms.…

Computation and Language · Computer Science 2021-06-04 Eric S. Tellez , Sabino Miranda-Jiménez , Mario Graff , Daniela Moctezuma , Oscar S. Siodia , Elio A. Villaseñor

In the dynamic realm of social media, diverse topics are discussed daily, transcending linguistic boundaries. However, the complexities of understanding and categorising this content across various languages remain an important challenge…

Computation and Language · Computer Science 2024-10-07 Dimosthenis Antypas , Asahi Ushio , Francesco Barbieri , Jose Camacho-Collados