中文
相关论文

相关论文: Regionalized models for Spanish language variation…

200 篇论文

The movements of ideas and content between locations and languages are unquestionably crucial concerns to researchers of the information age, and Twitter has emerged as a central, global platform on which hundreds of millions of people…

计算机与社会 · 计算机科学 2014-08-19 Mark Graham , Scott A. Hale , Devin Gaffney

The conventional natural language processing approaches are not accustomed to the social media text due to colloquial discourse and non-homogeneous characteristics. Significantly, the language identification in a multilingual document is…

计算与语言 · 计算机科学 2021-06-30 M Zeeshan Ansari , Tanvir Ahmad , M M Sufyan Beg , Asma Ikram

This paper describes a web-based corpus of global language use with a focus on how this corpus can be used for data-driven language mapping. First, the corpus provides a representation of where national varieties of major languages are used…

计算与语言 · 计算机科学 2020-04-03 Jonathan Dunn

Social media like Twitter provide a common platform to share and communicate personal experiences with other people. People often post their life experiences, local news, and events on social media to inform others. Many rescue agencies…

计算与语言 · 计算机科学 2021-08-25 Ashis Kumar Chanda

Pronunciation dictionaries are an important component in the process of speech forced alignment. The accuracy of these dictionaries has a strong effect on the aligned speech data since they help the mapping between orthographic…

计算与语言 · 计算机科学 2024-07-23 Simon Gonzalez

Sentiment Classification is a fundamental task in the field of Natural Language Processing, and has very important academic and commercial applications. It aims to automatically predict the degree of sentiment present in a text that…

计算与语言 · 计算机科学 2023-03-17 Lautaro Estienne , Matias Vera , Leonardo Rey Vega

Computer-mediated communication is driving fundamental changes in the nature of written language. We investigate these changes by statistical analysis of a dataset comprising 107 million Twitter messages (authored by 2.7 million unique user…

计算与语言 · 计算机科学 2014-11-25 Jacob Eisenstein , Brendan O'Connor , Noah A. Smith , Eric P. Xing

Twitter has become a pivotal platform for conducting information operations (IOs), particularly during high-stakes political events. In this study, we analyze over a million tweets about the 2024 U.S. presidential election to explore an…

社会与信息网络 · 计算机科学 2025-01-17 Bowen Yi

In applications involving conversational speech, data sparsity is a limiting factor in building a better language model. We propose a simple, language-independent method to quickly harvest large amounts of data from Twitter to supplement a…

计算与语言 · 计算机科学 2015-04-13 Aaron Jaech , Mari Ostendorf

Multilingual Transformer-based language models, usually pretrained on more than 100 languages, have been shown to achieve outstanding results in a wide range of cross-lingual transfer tasks. However, it remains unknown whether the…

计算与语言 · 计算机科学 2021-05-12 Laura Pérez-Mayos , Alba Táboas García , Simon Mille , Leo Wanner

This study presents an LLM-assisted annotation pipeline for the sociolinguistic and topical analysis of bilingual discourse in two typologically distinct contexts: Spanish-English and Spanish-Guaran\'i. Using large language models, we…

计算与语言 · 计算机科学 2025-12-04 Nemika Tyagi , Nelvin Licona Guevara , Olga Kellert

Investigating linguistic relationships on a global scale requires analyzing diverse features such as syntax, phonology and prosody, which evolve at varying rates influenced by internal diversification, language contact, and sociolinguistic…

计算与语言 · 计算机科学 2025-06-11 Tuukka Törö , Antti Suni , Juraj Šimko

Large language models (LLMs) offer new opportunities for scalable analysis of online discourse. Yet their use in multilingual social science research remains constrained by model size, cost and linguistic bias. We develop a lightweight,…

计算与语言 · 计算机科学 2025-12-30 Andrea Nasuto , Stefano Maria Iacus , Francisco Rowe , Devika Jain

Variations in languages across geographic regions or cultures are crucial to address to avoid biases in NLP systems designed for culturally sensitive tasks, such as hate speech detection or dialog with conversational agents. In languages…

计算与语言 · 计算机科学 2024-12-17 Javier A. Lopetegui , Arij Riabi , Djamé Seddah

Large Language Models (LLMs) exhibit inequalities with respect to various cultural contexts. Most prominent open-weights models are trained on Global North data and show prejudicial behavior towards other cultures. Moreover, there is a…

In contrast to much previous work that has focused on location classification of tweets restricted to a specific country, here we undertake the task in a broader context by classifying global tweets at the country level, which is so far…

信息检索 · 计算机科学 2017-04-26 Arkaitz Zubiaga , Alex Voss , Rob Procter , Maria Liakata , Bo Wang , Adam Tsakalidis

The content on the web is in a constant state of flux. New entities, issues, and ideas continuously emerge, while the semantics of the existing conversation topics gradually shift. In recent years, pre-trained language models like BERT…

计算与语言 · 计算机科学 2021-06-14 Spurthi Amba Hombaiah , Tao Chen , Mingyang Zhang , Michael Bendersky , Marc Najork

This work investigates the in-context learning abilities of pretrained large language models (LLMs) when instructed to translate text from a low-resource language into a high-resource language as part of an automated machine translation…

计算与语言 · 计算机科学 2024-10-28 Sara Court , Micha Elsner

We propose a simple yet effective text- based user geolocation model based on a neural network with one hidden layer, which achieves state of the art performance over three Twitter benchmark geolocation datasets, in addition to producing…

计算与语言 · 计算机科学 2017-04-28 Afshin Rahimi , Trevor Cohn , Timothy Baldwin

This work presents biomedical and clinical language models for Spanish by experimenting with different pretraining choices, such as masking at word and subword level, varying the vocabulary size and testing with domain data, looking for…