中文
相关论文

相关论文: Regionalized models for Spanish language variation…

200 篇论文

Training state-of-the-art offline handwriting recognition (HWR) models requires large labeled datasets, but unfortunately such datasets are not available in all languages and domains due to the high cost of manual labeling.We address this…

计算机视觉与模式识别 · 计算机科学 2018-08-07 Chris Tensmeyer , Curtis Wigington , Brian Davis , Seth Stewart , Tony Martinez , William Barrett

With the advent of digital technology, it is more common that committed crimes or legal disputes involve some form of speech recording where the identity of a speaker is questioned [1]. In face of this situation, the field of forensic…

声音 · 计算机科学 2021-09-15 Emmanuel Maqueda , Javier Alvarez-Jimenez , Carlos Mena , Ivan Meza

The present paper uses Twitter to analyze the current state of the worldwide, Spanish-language, independent publishing market. The main purposes are to determine whether certain Latin American Spanish-language independent publishers…

社会与信息网络 · 计算机科学 2020-08-04 Ana Gallego-Cuiñas , Esteban Romero-Frías , Wenceslao Arroyo-Machado

Several computational models have been developed to detect and analyze dialect variation in recent years. Most of these models assume a predefined set of geographical regions over which they detect and analyze dialectal variation. However,…

计算与语言 · 计算机科学 2019-10-17 Hang Jiang , Haoshen Hong , Yuxing Chen , Vivek Kulkarni

We present a system that allows users to train their own state-of-the-art paraphrastic sentence representations in a variety of languages. We also release trained models for English, Arabic, German, French, Spanish, Russian, Turkish, and…

计算与语言 · 计算机科学 2023-06-06 John Wieting , Kevin Gimpel , Graham Neubig , Taylor Berg-Kirkpatrick

Turkish is one of the most popular languages in the world. Wide us of this language on social media platforms such as Twitter, Instagram, or Tiktok and strategic position of the country in the world politics makes it appealing for the…

计算与语言 · 计算机科学 2023-12-01 Ali Najafi , Onur Varol

Much recent effort has been devoted to creating large-scale language models. Nowadays, the most prominent approaches are based on deep neural networks, such as BERT. However, they lack transparency and interpretability, and are often seen…

计算与语言 · 计算机科学 2023-06-08 Rodrigo Wilkens , Leonardo Zilio , Aline Villavicencio

Deep acoustic models represent linguistic information based on massive amounts of data. Unfortunately, for regional languages and dialects such resources are mostly not available. However, deep acoustic models might have learned linguistic…

计算与语言 · 计算机科学 2022-05-26 Martijn Bartelds , Martijn Wieling

Social media currently provide a window on our lives, making it possible to learn how people from different places, with different backgrounds, ages, and genders use language. In this work we exploit a newly-created Arabic dataset with…

计算与语言 · 计算机科学 2019-11-05 Muhammad Abdul-Mageed , Chiyu Zhang , Arun Rajendran , AbdelRahim Elmadany , Michael Przystupa , Lyle Ungar

Social media features substantial stylistic variation, raising new challenges for syntactic analysis of online writing. However, this variation is often aligned with author attributes such as age, gender, and geography, as well as more…

计算与语言 · 计算机科学 2018-04-23 Murali Raghu Babu Balusu , Taha Merghani , Jacob Eisenstein

This paper presents a new database collected from a bilingual speakers set (49), in two different languages: Spanish and Catalan. Phonetically there are significative differences between both languages. These differences have let us to…

音频与语音处理 · 电气工程与系统科学 2022-03-07 Antonio Satue-Villar , Marcos Faundez-Zanuy

While BERT is an effective method for learning monolingual sentence embeddings for semantic similarity and embedding based transfer learning (Reimers and Gurevych, 2019), BERT based cross-lingual sentence embeddings have yet to be explored.…

计算与语言 · 计算机科学 2022-03-09 Fangxiaoyu Feng , Yinfei Yang , Daniel Cer , Naveen Arivazhagan , Wei Wang

A considerable number of texts encountered daily are somehow connected with each other. For example, Wikipedia articles refer to other articles via hyperlinks, scientific papers relate to others via citations or (co)authors, while tweets…

计算与语言 · 计算机科学 2025-08-08 Albert Roethel , Maria Ganzha , Anna Wróblewska

Audio recordings of collaborative learning environments contain a constant presence of cross-talk and background noise. Dynamic speech recognition between Spanish and English is required in these environments. To eliminate the standard…

音频与语音处理 · 电气工程与系统科学 2022-02-23 Mario Esparza

The pervasiveness of mobile devices, which is increasing daily, is generating a vast amount of geo-located data allowing us to gain further insights into human behaviors. In particular, this new technology enables users to communicate…

物理与社会 · 物理学 2014-08-26 Maxime Lenormand , Antònia Tugores , Pere Colet , José J. Ramasco

Nowadays, digital news articles are widely available, published by various editors and often written in different languages. This large volume of diverse and unorganized information makes human reading very difficult or almost impossible.…

计算与语言 · 计算机科学 2020-04-20 Mathis Linger , Mhamed Hajaiej

Large language models are typically trained by treating text as a single global distribution, often resulting in geographically homogenized behavior. We study metadata conditioning as a lightweight approach for localization, pre-training 31…

计算与语言 · 计算机科学 2026-01-22 Anjishnu Mukherjee , Ziwei Zhu , Antonios Anastasopoulos

Hate speech detection deals with many language variants, slang, slurs, expression modalities, and cultural nuances. This outlines the importance of working with specific corpora, when addressing hate speech within the scope of Natural…

计算与语言 · 计算机科学 2024-10-17 Juan Manuel Pérez , Paula Miguel , Viviana Cotik

We present an evaluation of text simplification (TS) in Spanish for a production system, by means of two corpora focused in both complex-sentence and complex-word identification. We compare the most prevalent Spanish-specific readability…

计算与语言 · 计算机科学 2023-08-16 Adrian de Wynter , Anthony Hevia , Si-Qing Chen

Word embeddings provide an unsupervised way to understand differences in word usage between discursive communities. A number of recent papers have focused on identifying words that are used differently by two or more communities. But word…

计算与语言 · 计算机科学 2023-02-14 Thyge Enggaard , August Lohse , Morten Axel Pedersen , Sune Lehmann