中文
相关论文

相关论文: Learning about Spanish dialects through Twitter

200 篇论文

In the advent of a pervasive presence of location sharing services researchers gained an unprecedented access to the direct records of human activity in space and time. This paper analyses geo-located Twitter messages in order to uncover…

社会与信息网络 · 计算机科学 2014-03-03 Bartosz Hawelka , Izabela Sitko , Euro Beinat , Stanislav Sobolevsky , Pavlos Kazakopoulos , Carlo Ratti

Twitter is a popular public conversation platform with world-wide audience and diverse forms of connections between users. In this paper we introduce the concept of aggregated regional Twitter networks in order to characterize communication…

In this paper, we present Arap-Tweet, which is a large-scale and multi-dialectal corpus of Tweets from 11 regions and 16 countries in the Arab world representing the major Arabic dialectal varieties. To build this corpus, we collected data…

计算与语言 · 计算机科学 2018-08-24 Wajdi Zaghouani , Anis Charfi

Sentiment Classification is a fundamental task in the field of Natural Language Processing, and has very important academic and commercial applications. It aims to automatically predict the degree of sentiment present in a text that…

计算与语言 · 计算机科学 2023-03-17 Lautaro Estienne , Matias Vera , Leonardo Rey Vega

The geographical pattern of human dialects is a result of history. Here, we formulate a simple spatial model of language change which shows that the final result of this historical evolution may, to some extent, be predictable. The model…

物理与社会 · 物理学 2017-07-26 James Burridge

This research evidences the usefulness of open big data to map mobility patterns in a medium-sized city. Motivated by the novel analysis that big data allow worldwide and in large metropolitan areas, we developed a methodology aiming to…

社会与信息网络 · 计算机科学 2017-05-24 María Henar Salas-Olmedo , Carolina Rojas Quezada

Scaling properties of language are a useful tool for understanding generative processes in texts. We investigate the scaling relations in citywise Twitter corpora coming from the Metropolitan and Micropolitan Statistical Areas of the United…

物理与社会 · 物理学 2019-03-12 Eszter Bokányi , Dániel Kondor , Gábor Vattay

There are many Language Models for the English language according to its worldwide relevance. However, for the Spanish language, even if it is a widely spoken language, there are very few Spanish Language Models which result to be small and…

计算与语言 · 计算机科学 2021-10-26 Asier Gutiérrez-Fandiño , Jordi Armengol-Estapé , Aitor Gonzalez-Agirre , Marta Villegas

Mixed language data is one of the difficult yet less explored domains of natural language processing. Most research in fields like machine translation or sentiment analysis assume monolingual input. However, people who are capable of using…

神经与进化计算 · 计算机科学 2014-12-23 Joseph Chee Chang , Chu-Cheng Lin

This paper measures similarity both within and between 84 language varieties across nine languages. These corpora are drawn from digital sources (the web and tweets), allowing us to evaluate whether such geo-referenced corpora are reliable…

计算与语言 · 计算机科学 2021-04-06 Jonathan Dunn

Given the centrality of regions in social movements, politics and public administration we aim to quantitatively study inter- and intra-regional communication for the first time. This work uses social media posts to first identify…

社会与信息网络 · 计算机科学 2018-07-12 Rudy Arthur , Hywel T. P. Williams

We model and compute the probability distribution of the letters in random generated words in a language by using the theory of set partitions, Young tableaux and graph theoretical representation methods. This has been of interest for…

计算与语言 · 计算机科学 2014-07-24 Alberto Besana , Cristina Martínez

Pronunciation dictionaries are an important component in the process of speech forced alignment. The accuracy of these dictionaries has a strong effect on the aligned speech data since they help the mapping between orthographic…

计算与语言 · 计算机科学 2024-07-23 Simon Gonzalez

Text from social media provides a set of challenges that can cause traditional NLP approaches to fail. Informal language, spelling errors, abbreviations, and special characters are all commonplace in these posts, leading to a prohibitively…

机器学习 · 计算机科学 2016-05-18 Bhuwan Dhingra , Zhong Zhou , Dylan Fitzpatrick , Michael Muehl , William W. Cohen

Linguistic typology aims to capture structural and semantic variation across the world's languages. A large-scale typology could provide excellent guidance for multilingual Natural Language Processing (NLP), particularly for languages that…

In contrast to much previous work that has focused on location classification of tweets restricted to a specific country, here we undertake the task in a broader context by classifying global tweets at the country level, which is so far…

信息检索 · 计算机科学 2017-04-26 Arkaitz Zubiaga , Alex Voss , Rob Procter , Maria Liakata , Bo Wang , Adam Tsakalidis

Dialects represent a significant component of human culture and are found across all regions of the world. In Germany, more than 40% of the population speaks a regional dialect (Adler and Hansen, 2022). However, despite cultural importance,…

计算与语言 · 计算机科学 2025-09-18 Minh Duc Bui , Carolin Holtermann , Valentin Hofmann , Anne Lauscher , Katharina von der Wense

Variation in language is ubiquitous, particularly in newer forms of writing such as social media. Fortunately, variation is not random, it is often linked to social properties of the author. In this paper, we show how to exploit social…

计算与语言 · 计算机科学 2017-08-29 Yi Yang , Jacob Eisenstein

The meaning of a slang term can vary in different communities. However, slang semantic variation is not well understood and under-explored in the natural language processing of slang. One existing view argues that slang semantic variation…

计算与语言 · 计算机科学 2022-11-11 Zhewei Sun , Yang Xu

The Internet generates large volumes of data at a high rate, in particular, posts on social networks. Although social network data has numerous semantic adulterations, and is not intended to be a source of geo-spatial information, in the…