English
Related papers

Related papers: Learning about Spanish dialects through Twitter

200 papers

In the advent of a pervasive presence of location sharing services researchers gained an unprecedented access to the direct records of human activity in space and time. This paper analyses geo-located Twitter messages in order to uncover…

Social and Information Networks · Computer Science 2014-03-03 Bartosz Hawelka , Izabela Sitko , Euro Beinat , Stanislav Sobolevsky , Pavlos Kazakopoulos , Carlo Ratti

Twitter is a popular public conversation platform with world-wide audience and diverse forms of connections between users. In this paper we introduce the concept of aggregated regional Twitter networks in order to characterize communication…

In this paper, we present Arap-Tweet, which is a large-scale and multi-dialectal corpus of Tweets from 11 regions and 16 countries in the Arab world representing the major Arabic dialectal varieties. To build this corpus, we collected data…

Computation and Language · Computer Science 2018-08-24 Wajdi Zaghouani , Anis Charfi

Sentiment Classification is a fundamental task in the field of Natural Language Processing, and has very important academic and commercial applications. It aims to automatically predict the degree of sentiment present in a text that…

Computation and Language · Computer Science 2023-03-17 Lautaro Estienne , Matias Vera , Leonardo Rey Vega

The geographical pattern of human dialects is a result of history. Here, we formulate a simple spatial model of language change which shows that the final result of this historical evolution may, to some extent, be predictable. The model…

Physics and Society · Physics 2017-07-26 James Burridge

This research evidences the usefulness of open big data to map mobility patterns in a medium-sized city. Motivated by the novel analysis that big data allow worldwide and in large metropolitan areas, we developed a methodology aiming to…

Social and Information Networks · Computer Science 2017-05-24 María Henar Salas-Olmedo , Carolina Rojas Quezada

Scaling properties of language are a useful tool for understanding generative processes in texts. We investigate the scaling relations in citywise Twitter corpora coming from the Metropolitan and Micropolitan Statistical Areas of the United…

Physics and Society · Physics 2019-03-12 Eszter Bokányi , Dániel Kondor , Gábor Vattay

There are many Language Models for the English language according to its worldwide relevance. However, for the Spanish language, even if it is a widely spoken language, there are very few Spanish Language Models which result to be small and…

Computation and Language · Computer Science 2021-10-26 Asier Gutiérrez-Fandiño , Jordi Armengol-Estapé , Aitor Gonzalez-Agirre , Marta Villegas

Mixed language data is one of the difficult yet less explored domains of natural language processing. Most research in fields like machine translation or sentiment analysis assume monolingual input. However, people who are capable of using…

Neural and Evolutionary Computing · Computer Science 2014-12-23 Joseph Chee Chang , Chu-Cheng Lin

This paper measures similarity both within and between 84 language varieties across nine languages. These corpora are drawn from digital sources (the web and tweets), allowing us to evaluate whether such geo-referenced corpora are reliable…

Computation and Language · Computer Science 2021-04-06 Jonathan Dunn

Given the centrality of regions in social movements, politics and public administration we aim to quantitatively study inter- and intra-regional communication for the first time. This work uses social media posts to first identify…

Social and Information Networks · Computer Science 2018-07-12 Rudy Arthur , Hywel T. P. Williams

We model and compute the probability distribution of the letters in random generated words in a language by using the theory of set partitions, Young tableaux and graph theoretical representation methods. This has been of interest for…

Computation and Language · Computer Science 2014-07-24 Alberto Besana , Cristina Martínez

Pronunciation dictionaries are an important component in the process of speech forced alignment. The accuracy of these dictionaries has a strong effect on the aligned speech data since they help the mapping between orthographic…

Computation and Language · Computer Science 2024-07-23 Simon Gonzalez

Text from social media provides a set of challenges that can cause traditional NLP approaches to fail. Informal language, spelling errors, abbreviations, and special characters are all commonplace in these posts, leading to a prohibitively…

Machine Learning · Computer Science 2016-05-18 Bhuwan Dhingra , Zhong Zhou , Dylan Fitzpatrick , Michael Muehl , William W. Cohen

Linguistic typology aims to capture structural and semantic variation across the world's languages. A large-scale typology could provide excellent guidance for multilingual Natural Language Processing (NLP), particularly for languages that…

Computation and Language · Computer Science 2020-10-28 Edoardo Maria Ponti , Helen O'Horan , Yevgeni Berzak , Ivan Vulić , Roi Reichart , Thierry Poibeau , Ekaterina Shutova , Anna Korhonen

In contrast to much previous work that has focused on location classification of tweets restricted to a specific country, here we undertake the task in a broader context by classifying global tweets at the country level, which is so far…

Information Retrieval · Computer Science 2017-04-26 Arkaitz Zubiaga , Alex Voss , Rob Procter , Maria Liakata , Bo Wang , Adam Tsakalidis

Dialects represent a significant component of human culture and are found across all regions of the world. In Germany, more than 40% of the population speaks a regional dialect (Adler and Hansen, 2022). However, despite cultural importance,…

Computation and Language · Computer Science 2025-09-18 Minh Duc Bui , Carolin Holtermann , Valentin Hofmann , Anne Lauscher , Katharina von der Wense

Variation in language is ubiquitous, particularly in newer forms of writing such as social media. Fortunately, variation is not random, it is often linked to social properties of the author. In this paper, we show how to exploit social…

Computation and Language · Computer Science 2017-08-29 Yi Yang , Jacob Eisenstein

The meaning of a slang term can vary in different communities. However, slang semantic variation is not well understood and under-explored in the natural language processing of slang. One existing view argues that slang semantic variation…

Computation and Language · Computer Science 2022-11-11 Zhewei Sun , Yang Xu

The Internet generates large volumes of data at a high rate, in particular, posts on social networks. Although social network data has numerous semantic adulterations, and is not intended to be a source of geo-spatial information, in the…

Social and Information Networks · Computer Science 2022-09-08 Diana C. Pauca-Quispe , Cinthya Butron-Revilla , Ernesto Suarez-Lopez , Karla Aranibar-Tila , Jesus S. Aguilar-Ruiz