中文
相关论文

相关论文: Crowdsourcing Dialect Characterization through Twi…

200 篇论文

This paper maps the large-scale variation of the Spanish language by employing a corpus based on geographically tagged Twitter messages. Lexical dialects are extracted from an analysis of variants of tens of concepts. The resulting maps…

机器学习 · 统计学 2017-02-07 Bruno Gonçalves , David Sánchez

In the last few years, microblogging platforms such as Twitter have given rise to a deluge of textual data that can be used for the analysis of informal communication between millions of individuals. In this work, we propose an…

计算与语言 · 计算机科学 2021-11-17 Gonzalo Donoso , David Sanchez

Spanish is one of the most spoken languages in the globe, but not necessarily Spanish is written and spoken in the same way in different countries. Understanding local language variations can help to improve model performances on regional…

计算与语言 · 计算机科学 2022-12-13 Eric S. Tellez , Daniela Moctezuma , Sabino Miranda , Mario Graff , Guillermo Ruiz

The explosion in the availability of natural language data in the era of social media has given rise to a host of applications such as sentiment analysis and opinion mining. Simultaneously, the growing availability of precise geolocation…

计算与语言 · 计算机科学 2021-08-03 Olga Kellert , Nicholas H. Matlis

Large scale analysis and statistics of socio-technical systems that just a few short years ago would have required the use of consistent economic and human resources can nowadays be conveniently performed by mining the enormous amount of…

物理与社会 · 物理学 2013-04-23 Delia Mocanu , Andrea Baronchelli , Bruno Gonçalves , Nicola Perra , Alessandro Vespignani

The task of detecting regionalisms (expressions or words used in certain regions) has traditionally relied on the use of questionnaires and surveys, and has also heavily depended on the expertise and intuition of the surveyor. The irruption…

计算与语言 · 计算机科学 2019-07-11 Juan Manuel Pérez , Damián E. Aleman , Santiago N. Kalinowski , Agustín Gravano

This paper demonstrates a two-stage method for deriving insights from social media data relating to disinformation by applying a combination of geospatial classification and embedding-based language modelling across multiple languages. In…

Recent recollected data suggests that it is possible to automatically detect events that may negatively affect the most vulnerable parts of our society, by using any communication technology like social networks or messaging applications.…

This paper describes the system submitted to "Sentiment Analysis at SEPLN (TASS)-2019" shared task. The task includes sentiment analysis of Spanish tweets, where the tweets are in different dialects spoken in Spain, Peru, Costa Rica,…

计算与语言 · 计算机科学 2019-08-02 Avishek Garain , Sainik Kumar Mahata

Twitter has become a pivotal platform for conducting information operations (IOs), particularly during high-stakes political events. In this study, we analyze over a million tweets about the 2024 U.S. presidential election to explore an…

社会与信息网络 · 计算机科学 2025-01-17 Bowen Yi

Modern society habitually uses online social media services to publicly share observations, thoughts, opinions, and beliefs at any time and from any location. These geotagged social media posts may provide aggregate insights into people's…

社会与信息网络 · 计算机科学 2014-11-25 Derek Doran , Swapna Gokhale , Aldo Dagnino

In the dynamic realm of social media, diverse topics are discussed daily, transcending linguistic boundaries. However, the complexities of understanding and categorising this content across various languages remain an important challenge…

计算与语言 · 计算机科学 2024-10-07 Dimosthenis Antypas , Asahi Ushio , Francesco Barbieri , Jose Camacho-Collados

Much previous work characterizing language variation across Internet social groups has focused on the types of words used by these groups. We extend this type of study by employing BERT to characterize variation in the senses of words as…

计算与语言 · 计算机科学 2021-02-16 Li Lucy , David Bamman

Statistical linguistics has advanced considerably in recent decades as data has become available. This has allowed researchers to study how statistical properties of languages change over time. In this work, we use data from Twitter to…

Large language models are, by definition, based on language. In an effort to underscore the critical need for regional localized models, this paper examines primary differences between variants of written Spanish across Latin America and…

Recently, numerous approaches have emerged in the social sciences to exploit the opportunities made possible by the vast amounts of data generated by online social networks (OSNs). Having access to information about users on such a scale…

The socioeconomic background of people and how they use standard forms of language are not independent, as demonstrated in various sociolinguistic studies. However, the extent to which these correlations may be influenced by the mixing of…

物理与社会 · 物理学 2025-07-11 Thomas Louf , José J. Ramasco , David Sánchez , Márton Karsai

This study presents an LLM-assisted annotation pipeline for the sociolinguistic and topical analysis of bilingual discourse in two typologically distinct contexts: Spanish-English and Spanish-Guaran\'i. Using large language models, we…

计算与语言 · 计算机科学 2025-12-04 Nemika Tyagi , Nelvin Licona Guevara , Olga Kellert

Computer-mediated communication is driving fundamental changes in the nature of written language. We investigate these changes by statistical analysis of a dataset comprising 107 million Twitter messages (authored by 2.7 million unique user…

计算与语言 · 计算机科学 2014-11-25 Jacob Eisenstein , Brendan O'Connor , Noah A. Smith , Eric P. Xing

Sentiment Classification is a fundamental task in the field of Natural Language Processing, and has very important academic and commercial applications. It aims to automatically predict the degree of sentiment present in a text that…

计算与语言 · 计算机科学 2023-03-17 Lautaro Estienne , Matias Vera , Leonardo Rey Vega
‹ 上一页 1 2 3 10 下一页 ›