English
Related papers

Related papers: Regionalized models for Spanish language variation…

200 papers

This paper maps the large-scale variation of the Spanish language by employing a corpus based on geographically tagged Twitter messages. Lexical dialects are extracted from an analysis of variants of tens of concepts. The resulting maps…

Machine Learning · Statistics 2017-02-07 Bruno Gonçalves , David Sánchez

We perform a large-scale analysis of language diatopic variation using geotagged microblogging datasets. By collecting all Twitter messages written in Spanish over more than two years, we build a corpus from which a carefully selected list…

Physics and Society · Physics 2014-11-20 Bruno Gonçalves , David Sánchez

Large language models are, by definition, based on language. In an effort to underscore the critical need for regional localized models, this paper examines primary differences between variants of written Spanish across Latin America and…

In the last few years, microblogging platforms such as Twitter have given rise to a deluge of textual data that can be used for the analysis of informal communication between millions of individuals. In this work, we propose an…

Computation and Language · Computer Science 2021-11-17 Gonzalo Donoso , David Sanchez

The appearance of complex attention-based language models such as BERT, Roberta or GPT-3 has allowed to address highly complex tasks in a plethora of scenarios. However, when applied to specific domains, these models encounter considerable…

Computation and Language · Computer Science 2022-06-14 Javier Huertas-Tato , Alejandro Martin , David Camacho

The task of detecting regionalisms (expressions or words used in certain regions) has traditionally relied on the use of questionnaires and surveys, and has also heavily depended on the expertise and intuition of the surveyor. The irruption…

Computation and Language · Computer Science 2019-07-11 Juan Manuel Pérez , Damián E. Aleman , Santiago N. Kalinowski , Agustín Gravano

Since BERT appeared, Transformer language models and transfer learning have become state-of-the-art for Natural Language Understanding tasks. Recently, some works geared towards pre-training specially-crafted models for particular domains,…

Computation and Language · Computer Science 2022-05-05 Juan Manuel Pérez , Damián A. Furman , Laura Alonso Alemany , Franco Luque

The Spanish language is one of the top 5 spoken languages in the world. Nevertheless, finding resources to train or evaluate Spanish language models is not an easy task. In this paper we help bridge this gap by presenting a BERT-based…

Computation and Language · Computer Science 2023-08-08 José Cañete , Gabriel Chaperon , Rodrigo Fuentes , Jou-Hui Ho , Hojin Kang , Jorge Pérez

This study examines the extent to which Large Language Models (LLMs) capture geographic lexical variation in Spanish, a language that exhibits substantial regional variation. Treating LLMs as virtual informants, we probe their dialectal…

Computation and Language · Computer Science 2026-02-11 Yoshifumi Kawasaki

This paper demonstrates a two-stage method for deriving insights from social media data relating to disinformation by applying a combination of geospatial classification and embedding-based language modelling across multiple languages. In…

Computation and Language · Computer Science 2021-08-09 David Tuxworth , Dimosthenis Antypas , Luis Espinosa-Anke , Jose Camacho-Collados , Alun Preece , David Rogers

Statistical linguistics has advanced considerably in recent decades as data has become available. This has allowed researchers to study how statistical properties of languages change over time. In this work, we use data from Twitter to…

The explosion in the availability of natural language data in the era of social media has given rise to a host of applications such as sentiment analysis and opinion mining. Simultaneously, the growing availability of precise geolocation…

Computation and Language · Computer Science 2021-08-03 Olga Kellert , Nicholas H. Matlis

This paper addresses the critical gap in evaluating bias in multilingual Large Language Models (LLMs), with a specific focus on Spanish language within culturally-aware Latin American contexts. Despite widespread global deployment, current…

Computers and Society · Computer Science 2025-09-04 Melissa Robles , Catalina Bernal , Denniss Raigoso , Mateo Dulce Rubio

Suicidal ideation is a serious health problem affecting millions of people worldwide. Social networks provide information about these mental health problems through users' emotional expressions. We propose a multilingual model leveraging…

Computation and Language · Computer Science 2024-12-23 Rodolfo Zevallos , Annika Schoene , John E. Ortega

In recent years, large language models (LLMs) have demonstrated a high capacity for understanding and generating text in Spanish. However, with five hundred million native speakers, Spanish is not a homogeneous language but rather one rich…

Computation and Language · Computer Science 2025-05-22 Marina Mayor-Rocher , Cristina Pozo , Nina Melero , Gonzalo Martínez , María Grandury , Pedro Reviriego

Resources in high-resource languages have not been efficiently exploited in low-resource languages to solve language-dependent research problems. Spanish and French are considered high resource languages in which an adequate level of data…

Computation and Language · Computer Science 2023-12-13 Fatimah Alzamzami , Abdulmotaleb El Saddik

Language identification of social media text has been an interesting problem of study in recent years. Social media messages are predominantly in code mixed in non-English speaking states. Prior knowledge by pre-training contextual…

Computation and Language · Computer Science 2021-07-05 Mohd Zeeshan Ansari , M M Sufyan Beg , Tanvir Ahmad , Mohd Jazib Khan , Ghazali Wasim

Given the impact of language models on the field of Natural Language Processing, a number of Spanish encoder-only masked language models (aka BERTs) have been trained and released. These models were developed either within large projects…

Computation and Language · Computer Science 2023-09-25 Rodrigo Agerri , Eneko Agirre

Lexical ambiguity -- where a single wordform takes on distinct, context-dependent meanings -- serves as a useful tool to compare across different language models' (LMs') ability to form distinct, contextualized representations of the same…

Computation and Language · Computer Science 2025-03-14 Pamela D. Rivière , Anne L. Beatty-Martínez , Sean Trott

The popularity of social media has created problems such as hate speech and sexism. The identification and classification of sexism in social media are very relevant tasks, as they would allow building a healthier social environment.…

Computation and Language · Computer Science 2021-11-09 Angel Felipe Magnossão de Paula , Roberto Fray da Silva , Ipek Baris Schlicht
‹ Prev 1 2 3 10 Next ›