English
Related papers

Related papers: Detecting Linguistic Diversity on Social Media

200 papers

Accurate and timely population data are essential for disaster response and humanitarian planning, but traditional censuses often cannot capture rapid demographic changes. Social media data offer a promising alternative for dynamic…

Social media enables the rapid spread of many kinds of information, from memes to social movements. However, little is known about how information crosses linguistic boundaries. We apply causal inference techniques on the European Twitter…

Social and Information Networks · Computer Science 2023-04-11 Julia Mendelsohn , Sayan Ghosh , David Jurgens , Ceren Budak

Censuses around the world are key sources of data to guide government investments and public policies. However, these sources are very expensive to obtain and are collected relatively infrequently. Over the last decade, there has been…

Social and Information Networks · Computer Science 2020-05-19 Filipe N. Ribeiro , Fabrício Benevenuto , Emilio Zagheni

This study explores how different weather conditions influence public sentiment on social media, focusing on Twitter data from the UK. By considering climate and linguistic baselines, we improve the accuracy of weather-related sentiment…

Human-Computer Interaction · Computer Science 2024-07-11 James C. Young , Rudy Arthur , Hywel T. P. Williams

Social Internet content plays an increasingly critical role in many domains, including public health, disaster management, and politics. However, its utility is limited by missing geographic information; for example, fewer than 1.6% of…

Social and Information Networks · Computer Science 2013-11-19 Reid Priedhorsky , Aron Culotta , Sara Y. Del Valle

Computational measures of linguistic diversity help us understand the linguistic landscape using digital language data. The contribution of this paper is to calibrate measures of linguistic diversity using restrictions on international…

Computation and Language · Computer Science 2021-04-06 Jonathan Dunn , Tom Coupe , Benjamin Adams

The use of Large Language Models (LLMs) has proven to be a tool that could help in the automatic detection of sexism. Previous studies have shown that these models contain biases that do not accurately reflect reality, especially for…

Computation and Language · Computer Science 2025-08-26 Judith Tavarez-Rodríguez , Fernando Sánchez-Vega , A. Pastor López-Monroy

Technologies for abusive language detection are being developed and applied with little consideration of their potential biases. We examine racial bias in five different sets of Twitter data annotated for hate speech and abusive language.…

Computation and Language · Computer Science 2019-05-30 Thomas Davidson , Debasmita Bhattacharya , Ingmar Weber

We observe and report on a systematic relationship between population density and Twitter use. Number of tweets, number of users and population per unit area are related by power laws, with exponents greater than one, that are consistent…

Social and Information Networks · Computer Science 2017-11-28 Rudy Arthur , Hywel Williams

Statistical linguistics has advanced considerably in recent decades as data has become available. This has allowed researchers to study how statistical properties of languages change over time. In this work, we use data from Twitter to…

We use structural topic modeling to examine racial bias in data collected to train models to detect hate speech and abusive language in social media posts. We augment the abusive language dataset by adding an additional feature indicating…

Computation and Language · Computer Science 2020-05-28 Thomas Davidson , Debasmita Bhattacharya

The interest in demographic information retrieval based on text data has increased in the research community because applications have shown success in different sectors such as security, marketing, heath-care, and others. Recognition and…

Computation and Language · Computer Science 2021-07-07 Daniel Escobar-Grisales , Juan Camilo Vasquez-Correa , Juan Rafael Orozco-Arroyave

We investigate the predictive power behind the language of food on social media. We collect a corpus of over three million food-related posts from Twitter and demonstrate that many latent population characteristics can be directly predicted…

Computation and Language · Computer Science 2016-11-15 Daniel Fried , Mihai Surdeanu , Stephen Kobourov , Melanie Hingle , Dane Bell

Understanding the impact of digital platforms on user behavior presents foundational challenges, including issues related to polarization, misinformation dynamics, and variation in news consumption. Comparative analyses across platforms and…

Variation in language is ubiquitous, particularly in newer forms of writing such as social media. Fortunately, variation is not random, it is often linked to social properties of the author. In this paper, we show how to exploit social…

Computation and Language · Computer Science 2017-08-29 Yi Yang , Jacob Eisenstein

Statements on social media can be analysed to identify individuals who are experiencing red flag medical symptoms, allowing early detection of the spread of disease such as influenza. Since disease does not respect cultural borders and may…

Computation and Language · Computer Science 2019-10-11 Mattias Appelgren , Patrick Schrempf , Matúš Falis , Satoshi Ikeda , Alison Q O'Neil

Large Language Models (LLMs) have rapidly increased in size and apparent capabilities in the last three years, but their training data is largely English text. There is growing interest in multilingual LLMs, and various efforts are striving…

The prevalence of social media presents a growing opportunity to collect and analyse examples of English varieties. Whilst usage of these varieties was - and, in many cases, still is - used only in spoken contexts or hard-to-access private…

Computation and Language · Computer Science 2024-01-23 Nhi Pham , Lachlan Pham , Adam L. Meyers

Social media data provides propitious opportunities for public health research. However, studies suggest that disparities may exist in the representation of certain populations (e.g., people of lower socioeconomic status). To quantify and…

Computers and Society · Computer Science 2017-11-07 Nina Cesare , Christan Grant , Jared B. Hawkins , John S. Brownstein , Elaine O. Nsoesie

In today's global digital landscape, misinformation transcends linguistic boundaries, posing a significant challenge for moderation systems. Most approaches to misinformation detection are monolingual, focused on high-resource languages,…

Computation and Language · Computer Science 2025-04-01 Xinyu Wang , Wenbo Zhang , Sarah Rajtmajer