English
Related papers

Related papers: Dialect and Gender Bias in YouTube's Spanish Capti…

200 papers

Spanish is one of the most spoken languages in the globe, but not necessarily Spanish is written and spoken in the same way in different countries. Understanding local language variations can help to improve model performances on regional…

Computation and Language · Computer Science 2022-12-13 Eric S. Tellez , Daniela Moctezuma , Sabino Miranda , Mario Graff , Guillermo Ruiz

This paper maps the large-scale variation of the Spanish language by employing a corpus based on geographically tagged Twitter messages. Lexical dialects are extracted from an analysis of variants of tens of concepts. The resulting maps…

Machine Learning · Statistics 2017-02-07 Bruno Gonçalves , David Sánchez

In this study, we introduce YODAS (YouTube-Oriented Dataset for Audio and Speech), a large-scale, multilingual dataset comprising currently over 500k hours of speech data in more than 100 languages, sourced from both labeled and unlabeled…

Computation and Language · Computer Science 2024-06-04 Xinjian Li , Shinnosuke Takamichi , Takaaki Saeki , William Chen , Sayaka Shiota , Shinji Watanabe

The Internet's wealth of content, with up to 60% published in English, starkly contrasts the global population, where only 18.8% are English speakers, and just 5.1% consider it their native language, leading to disparities in online…

Audio and Speech Processing · Electrical Eng. & Systems 2024-01-11 Kevin Cai , Chonghua Liu , David M. Chan

Sentiment Classification is a fundamental task in the field of Natural Language Processing, and has very important academic and commercial applications. It aims to automatically predict the degree of sentiment present in a text that…

Computation and Language · Computer Science 2023-03-17 Lautaro Estienne , Matias Vera , Leonardo Rey Vega

Sexism is generally defined as prejudice and discrimination based on sex or gender, affecting every sector of society, from social institutions to relationships and individual behavior. Social media platforms amplify the impact of sexism by…

Computation and Language · Computer Science 2025-08-22 Laura De Grazia , Pol Pastells , Mauro Vázquez Chas , Desmond Elliott , Danae Sánchez Villegas , Mireia Farrús , Mariona Taulé

We construct the first ever multimodal sarcasm dataset for Spanish. The audiovisual dataset consists of sarcasm annotated text that is aligned with video and audio. The dataset represents two varieties of Spanish, a Latin American variety…

Computation and Language · Computer Science 2021-05-13 Khalid Alnajjar , Mika Hämäläinen

This study examines the extent to which Large Language Models (LLMs) capture geographic lexical variation in Spanish, a language that exhibits substantial regional variation. Treating LLMs as virtual informants, we probe their dialectal…

Computation and Language · Computer Science 2026-02-11 Yoshifumi Kawasaki

This paper presents VoxCeleb-ESP, a collection of pointers and timestamps to YouTube videos facilitating the creation of a novel speaker recognition dataset. VoxCeleb-ESP captures real-world scenarios, incorporating diverse speaking styles,…

We are 600 million Spanish speakers. We launched the #Somos600M Project because the diversity of the languages from LATAM, the Caribbean and Spain needs to be represented in Artificial Intelligence (AI) systems. Despite being the 7.5% of…

Computation and Language · Computer Science 2024-07-26 María Grandury

The Spanish language is one of the top 5 spoken languages in the world. Nevertheless, finding resources to train or evaluate Spanish language models is not an easy task. In this paper we help bridge this gap by presenting a BERT-based…

Computation and Language · Computer Science 2023-08-08 José Cañete , Gabriel Chaperon , Rodrigo Fuentes , Jou-Hui Ho , Hojin Kang , Jorge Pérez

Machine Translation (MT) continues to make significant strides in quality and is increasingly adopted on a larger scale. Consequently, analyses have been redirected to more nuanced aspects, intricate phenomena, as well as potential risks…

Computation and Language · Computer Science 2024-03-28 Silvia Alma Piazzolla , Beatrice Savoldi , Luisa Bentivogli

Large language models are, by definition, based on language. In an effort to underscore the critical need for regional localized models, this paper examines primary differences between variants of written Spanish across Latin America and…

Machine learning for sign languages is bottlenecked by data. In this paper, we present YouTube-ASL, a large-scale, open-domain corpus of American Sign Language (ASL) videos and accompanying English captions drawn from YouTube. With ~1000…

Computation and Language · Computer Science 2023-10-30 David Uthus , Garrett Tanzer , Manfred Georg

This work addresses the development of a database for the automatic assessment of text-to-speech (TTS) systems in Spanish, aiming to improve the accuracy of naturalness prediction models. The dataset consists of 4,326 audio samples from 52…

Sound · Computer Science 2025-07-03 Alejandro Sosa Welford , Leonardo Pepino

Cultural and language factors significantly influence counseling, but Natural Language Processing research has not yet examined whether the findings of conversational analysis for counseling conducted in English apply to other languages.…

Computation and Language · Computer Science 2025-02-13 Aylin Gunal , Bowen Yi , John Piette , Rada Mihalcea , Verónica Pérez-Rosas

Pronunciation dictionaries are an important component in the process of speech forced alignment. The accuracy of these dictionaries has a strong effect on the aligned speech data since they help the mapping between orthographic…

Computation and Language · Computer Science 2024-07-23 Simon Gonzalez

Abusive language is a massive problem in online social platforms. Existing abusive language detection techniques are particularly ill-suited to comments containing heterogeneous abusive language patterns, i.e., both abusive and non-abusive…

Computation and Language · Computer Science 2021-05-25 Hongyu Gong , Alberto Valido , Katherine M. Ingram , Giulia Fanti , Suma Bhat , Dorothy L. Espelage

This paper investigates the use of automatically collected web audio data for the task of spoken language recognition. We generate semi-random search phrases from language-specific Wikipedia data that are then used to retrieve videos from…

Audio and Speech Processing · Electrical Eng. & Systems 2020-11-30 Jörgen Valk , Tanel Alumäe

The advent of social media has transformed communication, enabling individuals to share their experiences, seek support, and participate in diverse discussions. While extensive research has focused on identifying harmful content like hate…

Computation and Language · Computer Science 2025-02-17 Moein Shahiki Tash , Luis Ramos , Zahra Ahani , Raul Monroy , Olga kolesnikova , Hiram Calvo , Grigori Sidorov
‹ Prev 1 2 3 10 Next ›