English
Related papers

Related papers: Mapping Languages and Demographics with Georeferen…

200 papers

While extremely useful (e.g., for COVID-19 forecasting and policy-making, urban mobility analysis and marketing, and obtaining business insights), location data collected from mobile devices often contain data from a biased population…

Machine Learning · Computer Science 2024-02-20 Sepanta Zeighami , Cyrus Shahabi

High-resolution human settlement maps provide detailed delineations of where people live and are vital for scientific and practical purposes, such as rapid disaster response, allocation of humanitarian resources, and international…

Social and Information Networks · Computer Science 2024-04-23 Vedran Sekara , Andrea Martini , Manuel Garcia-Herranz , Do-Hyung Kim

The increasing tendency to collect large and uncurated datasets to train vision-and-language models has raised concerns about fair representations. It is known that even small but manually annotated datasets, such as MSCOCO, are affected by…

Computer Vision and Pattern Recognition · Computer Science 2023-04-07 Noa Garcia , Yusuke Hirota , Yankun Wu , Yuta Nakashima

Word embeddings are usually derived from corpora containing text from many individuals, thus leading to general purpose representations rather than individually personalized representations. While personalized embeddings can be useful to…

Computation and Language · Computer Science 2020-11-22 Charles Welch , Jonathan K. Kummerfeld , Verónica Pérez-Rosas , Rada Mihalcea

The socioeconomic background of people and how they use standard forms of language are not independent, as demonstrated in various sociolinguistic studies. However, the extent to which these correlations may be influenced by the mixing of…

Physics and Society · Physics 2025-07-11 Thomas Louf , José J. Ramasco , David Sánchez , Márton Karsai

We introduce a multilingual extension of the HOLISTICBIAS dataset, the largest English template-based taxonomy of textual people references: MULTILINGUALHOLISTICBIAS. This extension consists of 20,459 sentences in 50 languages distributed…

Media framing is the study of strategically selecting and presenting specific aspects of political issues to shape public opinion. Despite its relevance to almost all societies around the world, research has been limited due to the lack of…

Computation and Language · Computer Science 2024-04-03 Syeda Sabrina Akter , Antonios Anastasopoulos

Massive web-crawled image-text datasets lay the foundation for recent progress in multimodal learning. These datasets are designed with the goal of training a model to do well on standard computer vision benchmarks, many of which, however,…

Computer Vision and Pattern Recognition · Computer Science 2025-09-16 Thao Nguyen , Matthew Wallingford , Sebastin Santy , Wei-Chiu Ma , Sewoong Oh , Ludwig Schmidt , Pang Wei Koh , Ranjay Krishna

This paper demonstrates a two-stage method for deriving insights from social media data relating to disinformation by applying a combination of geospatial classification and embedding-based language modelling across multiple languages. In…

Computation and Language · Computer Science 2021-08-09 David Tuxworth , Dimosthenis Antypas , Luis Espinosa-Anke , Jose Camacho-Collados , Alun Preece , David Rogers

We investigate in this paper how distributions of occupations with respect to gender is reflected in pre-trained language models. Such distributions are not always aligned to normative ideals, nor do they necessarily reflect a descriptive…

Computation and Language · Computer Science 2023-04-13 Samia Touileb , Lilja Øvrelid , Erik Velldal

Social media outlets such as Twitter constitute valuable data sources for understanding human activities in the virtual world from a geographic perspective. This paper examines spatial distribution of tweets and densities within cities. The…

Physics and Society · Physics 2020-09-04 Bin Jiang , Ding Ma , Junjun Yin , Mats Sandberg

A neural language model trained on a text corpus can be used to induce distributed representations of words, such that similar words end up with similar representations. If the corpus is multilingual, the same model can be used to learn…

Computation and Language · Computer Science 2019-01-10 Johannes Bjerva , Robert Östling , Maria Han Veiga , Jörg Tiedemann , Isabelle Augenstein

Modern models for common NLP tasks often employ machine learning techniques and train on journalistic, social media, or other culturally-derived text. These have recently been scrutinized for racial and gender biases, rooting from inherent…

Computation and Language · Computer Science 2026-01-27 Scott Friedman , Sonja Schmer-Galunder , Anthony Chen , Jeffrey Rye

Understanding human migration is of great interest to demographers and social scientists. User generated digital data has made it easier to study such patterns at a global scale. Geo coded Twitter data, in particular, has been shown to be a…

Social and Information Networks · Computer Science 2017-02-17 Hieu Nguyen , Kiran Garimella

Word choice is dependent on the cultural context of writers and their subjects. Different words are used to describe similar actions, objects, and features based on factors such as class, race, gender, geography and political affinity.…

Computation and Language · Computer Science 2018-07-02 Taylor Arnold , Lauren Tilton

The social connections people form online affect the quality of information they receive and their online experience. Although a host of socioeconomic and cognitive factors were implicated in the formation of offline social ties, few of…

Computers and Society · Computer Science 2016-03-11 Kristina Lerman , Megha Arora , Luciano Gallegos , Ponnurangam Kumaraguru , David Garcia

Personality and demographics are important variables in social sciences, while in NLP they can aid in interpretability and removal of societal biases. However, datasets with both personality and demographic labels are scarce. To address…

Computation and Language · Computer Science 2021-06-09 Matej Gjurković , Mladen Karan , Iva Vukojević , Mihaela Bošnjak , Jan Šnajder

Social media contains useful information about people and the society that could help advance research in many different areas (e.g. by applying opinion mining, emotion/sentiment analysis, and statistical analysis) such as business and…

Computation and Language · Computer Science 2022-05-16 Zahra Movahedi Nia , Ali Ahmadi , Bruce Mellado , Jianhong Wu , James Orbinski , Ali Agary , Jude Dzevela Kong

A vast amount of geographic information exists in natural language texts, such as tweets and news. Extracting geographic information from texts is called Geoparsing, which includes two subtasks: toponym recognition and toponym…

Information Retrieval · Computer Science 2022-09-20 Xuke Hu , Yeran Sun , Jens Kersten , Zhiyong Zhou , Friederike Klan , Hongchao Fan

The increasing prevalence of location-sharing features on social media has enabled researchers to ground computational social science research using geolocated data, affording opportunities to study human mobility, the impact of real-world…

Social and Information Networks · Computer Science 2023-02-15 Julie Jiang , Jesse Thomason , Francesco Barbieri , Emilio Ferrara