English
Related papers

Related papers: Dialectometric analysis of language variation in T…

200 papers

Modelling and forecasting real-life human behaviour using online social media is an active endeavour of interest in politics, government, academia, and industry. Since its creation in 2006, Twitter has been proposed as a potential…

Social and Information Networks · Computer Science 2023-08-15 Alejandro Vigna-Gómez , Javier Murillo , Manelik Ramirez , Alberto Borbolla , Ian Márquez , Prasun K. Ray

We investigate the predictive power behind the language of food on social media. We collect a corpus of over three million food-related posts from Twitter and demonstrate that many latent population characteristics can be directly predicted…

Computation and Language · Computer Science 2016-11-15 Daniel Fried , Mihai Surdeanu , Stephen Kobourov , Melanie Hingle , Dane Bell

In this paper, we present Arap-Tweet, which is a large-scale and multi-dialectal corpus of Tweets from 11 regions and 16 countries in the Arab world representing the major Arabic dialectal varieties. To build this corpus, we collected data…

Computation and Language · Computer Science 2018-08-24 Wajdi Zaghouani , Anis Charfi

In today's digital landscape, the proliferation of conspiracy theories within the disinformation ecosystem of online platforms represents a growing concern. This paper delves into the complexities of this phenomenon. We conducted a…

Social and Information Networks · Computer Science 2024-05-22 Alessandra Recordare , Guglielmo Cola , Tiziano Fagni , Maurizio Tesconi

Geolocalization of social media content is the task of determining the geographical location of a user based on textual data, that may show linguistic variations and informal language. In this project, we address the GeoLingIt challenge of…

Computation and Language · Computer Science 2024-07-24 Davide Savarro , Davide Zago , Stefano Zoia

Semantic sentence embeddings are usually supervisedly built minimizing distances between pairs of embeddings of sentences labelled as semantically similar by annotators. Since big labelled datasets are rare, in particular for non-English…

Computation and Language · Computer Science 2021-10-06 Marco Di Giovanni , Marco Brambilla

Micro-blogging services can track users' geo-locations when users check-in their places or use geo-tagging which implicitly reveals locations. This "geo tracking" can help to find topics triggered by some events in certain regions. However,…

Information Retrieval · Computer Science 2016-07-21 Siwei Qiang , Yongkun Wang , Yaohui Jin

The content on the web is in a constant state of flux. New entities, issues, and ideas continuously emerge, while the semantics of the existing conversation topics gradually shift. In recent years, pre-trained language models like BERT…

Computation and Language · Computer Science 2021-06-14 Spurthi Amba Hombaiah , Tao Chen , Mingyang Zhang , Michael Bendersky , Marc Najork

The geolocation of online information is an essential component in any geospatial application. While most of the previous work on geolocation has focused on Twitter, in this paper we quantify and compare the performance of text-based…

Computation and Language · Computer Science 2018-11-20 Konstantinos Pappas , Mahmoud Azab , Rada Mihalcea

In this paper we take into account both social and linguistic aspects to perform demographic analysis by processing a large amount of tweets in Basque language. The study of demographic characteristics and social relationships are…

Computers and Society · Computer Science 2021-09-09 J. Fernandez de Landa , R. Agerri

We propose a method for embedding two-dimensional locations in a continuous vector space using a neural network-based model incorporating mixtures of Gaussian distributions, presenting two model variants for text-based geolocation and…

Computation and Language · Computer Science 2017-08-16 Afshin Rahimi , Timothy Baldwin , Trevor Cohn

Cluster analysis is a field of data analysis that extracts underlying patterns in data. One application of cluster analysis is in text-mining, the analysis of large collections of text to find similarities between documents. We used a…

Machine Learning · Statistics 2014-08-26 Daniel Godfrey , Caley Johns , Carl Meyer , Shaina Race , Carol Sadek

This study aims to comprehend linguistic and socio-demographic features, encompassing English language styles, conveyed sentiments, and lexical diversity within spatial online social media review data. To this end, we undertake a case study…

Computation and Language · Computer Science 2023-11-14 Salim Sazzed

Researchers since at least Darwin have debated whether and to what extent emotions are universal or culture-dependent. However, previous studies have primarily focused on facial expressions and on a limited set of emotions. Given that…

Computation and Language · Computer Science 2013-04-30 Eugene Yuta Bann , Joanna J. Bryson

Quantifying the degree of spatial dependence for linguistic variables is a key task for analyzing dialectal variation. However, existing approaches have important drawbacks. First, they are based on parametric models of dependence, which…

Computation and Language · Computer Science 2016-08-30 Dong Nguyen , Jacob Eisenstein

Identifying controversial topics is not only interesting from a social point of view, it also enables the application of methods to avoid the information segregation, creating better discussion contexts and reaching agreements in the best…

Information Retrieval · Computer Science 2020-01-28 Juan Manuel Ortiz de Zarate , Esteban Feuerstein

Social Media have been extensively used for commercial and political communication, besides their initial scope of providing an easy-to-use outlet to produce and consume user-generated content. Besides being a popular medium, Social Media…

Social and Information Networks · Computer Science 2022-10-17 Kostas Karpouzis , Stavros Kaperonis , Yannis Skarpelos

In recent years, multimodal natural language processing, aimed at learning from diverse data types, has garnered significant attention. However, there needs to be more clarity when it comes to analysing multimodal tasks in multi-lingual…

Computation and Language · Computer Science 2024-06-13 Gaurish Thakkar , Sherzod Hakimov , Marko Tadić

This paper introduces a large collection of time series data derived from Twitter, postprocessed using word embedding techniques, as well as specialized fine-tuned language models. This data comprises the past five years and captures…

Computation and Language · Computer Science 2023-08-07 Daniel Loureiro , Kiamehr Rezaee , Talayeh Riahi , Francesco Barbieri , Leonardo Neves , Luis Espinosa Anke , Jose Camacho-Collados

In the advent of a pervasive presence of location sharing services researchers gained an unprecedented access to the direct records of human activity in space and time. This paper analyses geo-located Twitter messages in order to uncover…

Social and Information Networks · Computer Science 2014-03-03 Bartosz Hawelka , Izabela Sitko , Euro Beinat , Stanislav Sobolevsky , Pavlos Kazakopoulos , Carlo Ratti
‹ Prev 1 3 4 5 6 7 10 Next ›