English
Related papers

Related papers: Regionalized models for Spanish language variation…

200 papers

Training state-of-the-art offline handwriting recognition (HWR) models requires large labeled datasets, but unfortunately such datasets are not available in all languages and domains due to the high cost of manual labeling.We address this…

Computer Vision and Pattern Recognition · Computer Science 2018-08-07 Chris Tensmeyer , Curtis Wigington , Brian Davis , Seth Stewart , Tony Martinez , William Barrett

With the advent of digital technology, it is more common that committed crimes or legal disputes involve some form of speech recording where the identity of a speaker is questioned [1]. In face of this situation, the field of forensic…

Sound · Computer Science 2021-09-15 Emmanuel Maqueda , Javier Alvarez-Jimenez , Carlos Mena , Ivan Meza

The present paper uses Twitter to analyze the current state of the worldwide, Spanish-language, independent publishing market. The main purposes are to determine whether certain Latin American Spanish-language independent publishers…

Social and Information Networks · Computer Science 2020-08-04 Ana Gallego-Cuiñas , Esteban Romero-Frías , Wenceslao Arroyo-Machado

Several computational models have been developed to detect and analyze dialect variation in recent years. Most of these models assume a predefined set of geographical regions over which they detect and analyze dialectal variation. However,…

Computation and Language · Computer Science 2019-10-17 Hang Jiang , Haoshen Hong , Yuxing Chen , Vivek Kulkarni

We present a system that allows users to train their own state-of-the-art paraphrastic sentence representations in a variety of languages. We also release trained models for English, Arabic, German, French, Spanish, Russian, Turkish, and…

Computation and Language · Computer Science 2023-06-06 John Wieting , Kevin Gimpel , Graham Neubig , Taylor Berg-Kirkpatrick

Turkish is one of the most popular languages in the world. Wide us of this language on social media platforms such as Twitter, Instagram, or Tiktok and strategic position of the country in the world politics makes it appealing for the…

Computation and Language · Computer Science 2023-12-01 Ali Najafi , Onur Varol

Much recent effort has been devoted to creating large-scale language models. Nowadays, the most prominent approaches are based on deep neural networks, such as BERT. However, they lack transparency and interpretability, and are often seen…

Computation and Language · Computer Science 2023-06-08 Rodrigo Wilkens , Leonardo Zilio , Aline Villavicencio

Deep acoustic models represent linguistic information based on massive amounts of data. Unfortunately, for regional languages and dialects such resources are mostly not available. However, deep acoustic models might have learned linguistic…

Computation and Language · Computer Science 2022-05-26 Martijn Bartelds , Martijn Wieling

Social media currently provide a window on our lives, making it possible to learn how people from different places, with different backgrounds, ages, and genders use language. In this work we exploit a newly-created Arabic dataset with…

Computation and Language · Computer Science 2019-11-05 Muhammad Abdul-Mageed , Chiyu Zhang , Arun Rajendran , AbdelRahim Elmadany , Michael Przystupa , Lyle Ungar

Social media features substantial stylistic variation, raising new challenges for syntactic analysis of online writing. However, this variation is often aligned with author attributes such as age, gender, and geography, as well as more…

Computation and Language · Computer Science 2018-04-23 Murali Raghu Babu Balusu , Taha Merghani , Jacob Eisenstein

This paper presents a new database collected from a bilingual speakers set (49), in two different languages: Spanish and Catalan. Phonetically there are significative differences between both languages. These differences have let us to…

Audio and Speech Processing · Electrical Eng. & Systems 2022-03-07 Antonio Satue-Villar , Marcos Faundez-Zanuy

While BERT is an effective method for learning monolingual sentence embeddings for semantic similarity and embedding based transfer learning (Reimers and Gurevych, 2019), BERT based cross-lingual sentence embeddings have yet to be explored.…

Computation and Language · Computer Science 2022-03-09 Fangxiaoyu Feng , Yinfei Yang , Daniel Cer , Naveen Arivazhagan , Wei Wang

A considerable number of texts encountered daily are somehow connected with each other. For example, Wikipedia articles refer to other articles via hyperlinks, scientific papers relate to others via citations or (co)authors, while tweets…

Computation and Language · Computer Science 2025-08-08 Albert Roethel , Maria Ganzha , Anna Wróblewska

Audio recordings of collaborative learning environments contain a constant presence of cross-talk and background noise. Dynamic speech recognition between Spanish and English is required in these environments. To eliminate the standard…

Audio and Speech Processing · Electrical Eng. & Systems 2022-02-23 Mario Esparza

The pervasiveness of mobile devices, which is increasing daily, is generating a vast amount of geo-located data allowing us to gain further insights into human behaviors. In particular, this new technology enables users to communicate…

Physics and Society · Physics 2014-08-26 Maxime Lenormand , Antònia Tugores , Pere Colet , José J. Ramasco

Nowadays, digital news articles are widely available, published by various editors and often written in different languages. This large volume of diverse and unorganized information makes human reading very difficult or almost impossible.…

Computation and Language · Computer Science 2020-04-20 Mathis Linger , Mhamed Hajaiej

Large language models are typically trained by treating text as a single global distribution, often resulting in geographically homogenized behavior. We study metadata conditioning as a lightweight approach for localization, pre-training 31…

Computation and Language · Computer Science 2026-01-22 Anjishnu Mukherjee , Ziwei Zhu , Antonios Anastasopoulos

Hate speech detection deals with many language variants, slang, slurs, expression modalities, and cultural nuances. This outlines the importance of working with specific corpora, when addressing hate speech within the scope of Natural…

Computation and Language · Computer Science 2024-10-17 Juan Manuel Pérez , Paula Miguel , Viviana Cotik

We present an evaluation of text simplification (TS) in Spanish for a production system, by means of two corpora focused in both complex-sentence and complex-word identification. We compare the most prevalent Spanish-specific readability…

Computation and Language · Computer Science 2023-08-16 Adrian de Wynter , Anthony Hevia , Si-Qing Chen

Word embeddings provide an unsupervised way to understand differences in word usage between discursive communities. A number of recent papers have focused on identifying words that are used differently by two or more communities. But word…

Computation and Language · Computer Science 2023-02-14 Thyge Enggaard , August Lohse , Morten Axel Pedersen , Sune Lehmann