中文
相关论文

相关论文: Regionalized models for Spanish language variation…

200 篇论文

This paper maps the large-scale variation of the Spanish language by employing a corpus based on geographically tagged Twitter messages. Lexical dialects are extracted from an analysis of variants of tens of concepts. The resulting maps…

机器学习 · 统计学 2017-02-07 Bruno Gonçalves , David Sánchez

We perform a large-scale analysis of language diatopic variation using geotagged microblogging datasets. By collecting all Twitter messages written in Spanish over more than two years, we build a corpus from which a carefully selected list…

物理与社会 · 物理学 2014-11-20 Bruno Gonçalves , David Sánchez

Large language models are, by definition, based on language. In an effort to underscore the critical need for regional localized models, this paper examines primary differences between variants of written Spanish across Latin America and…

In the last few years, microblogging platforms such as Twitter have given rise to a deluge of textual data that can be used for the analysis of informal communication between millions of individuals. In this work, we propose an…

计算与语言 · 计算机科学 2021-11-17 Gonzalo Donoso , David Sanchez

The appearance of complex attention-based language models such as BERT, Roberta or GPT-3 has allowed to address highly complex tasks in a plethora of scenarios. However, when applied to specific domains, these models encounter considerable…

计算与语言 · 计算机科学 2022-06-14 Javier Huertas-Tato , Alejandro Martin , David Camacho

The task of detecting regionalisms (expressions or words used in certain regions) has traditionally relied on the use of questionnaires and surveys, and has also heavily depended on the expertise and intuition of the surveyor. The irruption…

计算与语言 · 计算机科学 2019-07-11 Juan Manuel Pérez , Damián E. Aleman , Santiago N. Kalinowski , Agustín Gravano

Since BERT appeared, Transformer language models and transfer learning have become state-of-the-art for Natural Language Understanding tasks. Recently, some works geared towards pre-training specially-crafted models for particular domains,…

计算与语言 · 计算机科学 2022-05-05 Juan Manuel Pérez , Damián A. Furman , Laura Alonso Alemany , Franco Luque

The Spanish language is one of the top 5 spoken languages in the world. Nevertheless, finding resources to train or evaluate Spanish language models is not an easy task. In this paper we help bridge this gap by presenting a BERT-based…

计算与语言 · 计算机科学 2023-08-08 José Cañete , Gabriel Chaperon , Rodrigo Fuentes , Jou-Hui Ho , Hojin Kang , Jorge Pérez

This study examines the extent to which Large Language Models (LLMs) capture geographic lexical variation in Spanish, a language that exhibits substantial regional variation. Treating LLMs as virtual informants, we probe their dialectal…

计算与语言 · 计算机科学 2026-02-11 Yoshifumi Kawasaki

This paper demonstrates a two-stage method for deriving insights from social media data relating to disinformation by applying a combination of geospatial classification and embedding-based language modelling across multiple languages. In…

Statistical linguistics has advanced considerably in recent decades as data has become available. This has allowed researchers to study how statistical properties of languages change over time. In this work, we use data from Twitter to…

The explosion in the availability of natural language data in the era of social media has given rise to a host of applications such as sentiment analysis and opinion mining. Simultaneously, the growing availability of precise geolocation…

计算与语言 · 计算机科学 2021-08-03 Olga Kellert , Nicholas H. Matlis

This paper addresses the critical gap in evaluating bias in multilingual Large Language Models (LLMs), with a specific focus on Spanish language within culturally-aware Latin American contexts. Despite widespread global deployment, current…

计算机与社会 · 计算机科学 2025-09-04 Melissa Robles , Catalina Bernal , Denniss Raigoso , Mateo Dulce Rubio

Suicidal ideation is a serious health problem affecting millions of people worldwide. Social networks provide information about these mental health problems through users' emotional expressions. We propose a multilingual model leveraging…

计算与语言 · 计算机科学 2024-12-23 Rodolfo Zevallos , Annika Schoene , John E. Ortega

In recent years, large language models (LLMs) have demonstrated a high capacity for understanding and generating text in Spanish. However, with five hundred million native speakers, Spanish is not a homogeneous language but rather one rich…

计算与语言 · 计算机科学 2025-05-22 Marina Mayor-Rocher , Cristina Pozo , Nina Melero , Gonzalo Martínez , María Grandury , Pedro Reviriego

Resources in high-resource languages have not been efficiently exploited in low-resource languages to solve language-dependent research problems. Spanish and French are considered high resource languages in which an adequate level of data…

计算与语言 · 计算机科学 2023-12-13 Fatimah Alzamzami , Abdulmotaleb El Saddik

Language identification of social media text has been an interesting problem of study in recent years. Social media messages are predominantly in code mixed in non-English speaking states. Prior knowledge by pre-training contextual…

计算与语言 · 计算机科学 2021-07-05 Mohd Zeeshan Ansari , M M Sufyan Beg , Tanvir Ahmad , Mohd Jazib Khan , Ghazali Wasim

Given the impact of language models on the field of Natural Language Processing, a number of Spanish encoder-only masked language models (aka BERTs) have been trained and released. These models were developed either within large projects…

计算与语言 · 计算机科学 2023-09-25 Rodrigo Agerri , Eneko Agirre

Lexical ambiguity -- where a single wordform takes on distinct, context-dependent meanings -- serves as a useful tool to compare across different language models' (LMs') ability to form distinct, contextualized representations of the same…

计算与语言 · 计算机科学 2025-03-14 Pamela D. Rivière , Anne L. Beatty-Martínez , Sean Trott

The popularity of social media has created problems such as hate speech and sexism. The identification and classification of sexism in social media are very relevant tasks, as they would allow building a healthier social environment.…

计算与语言 · 计算机科学 2021-11-09 Angel Felipe Magnossão de Paula , Roberto Fray da Silva , Ipek Baris Schlicht
‹ 上一页 1 2 3 10 下一页 ›