中文
相关论文

相关论文: Mapping Languages and Demographics with Georeferen…

200 篇论文

An important aspect of urban planning is understanding crowd levels at various locations, which typically require the use of physical sensors. Such sensors are potentially costly and time consuming to implement on a large scale. To address…

社会与信息网络 · 计算机科学 2020-12-08 Jerome Heng , Junhua Liu , Kwan Hui Lim

What is a population? This review considers how a population may be defined in terms of understanding the structure of the underlying genetics of the individuals involved. The main approach is to consider statistically identifiable groups…

种群与进化 · 定量生物学 2013-06-05 Daniel John Lawson

Personal electronic devices including smartphones give access to behavioural signals that can be used to learn about the characteristics and preferences of individuals. In this study, we explore the connection between demographic and…

计算机与社会 · 计算机科学 2018-11-26 Kyriaki Kalimeri , Mariano G. Beiro , Matteo Delfino , Robert Raleigh , Ciro Cattuto

In order to demonstrate why it is important to correctly account for the (serial dependent) structure of temporal data, we document an apparently spectacular relationship between population size and lexical diversity: for five out of seven…

计算与语言 · 计算机科学 2016-04-27 Alexander Koplenig , Carolin Mueller-Spitzer

Large Language Models (LLMs) have seen widespread deployment in various real-world applications. Understanding these biases is crucial to comprehend the potential downstream consequences when using LLMs to make decisions, particularly for…

计算与语言 · 计算机科学 2024-01-10 Abel Salinas , Parth Vipul Shah , Yuzhong Huang , Robert McCormack , Fred Morstatter

Geo-tagged Twitter data has been used recently to infer insights on the human aspects of social media. Insights related to demographics, spatial distribution of cultural activities, space-time travel trajectories for humans as well as…

社会与信息网络 · 计算机科学 2019-07-30 Wei Lun Lim , Chiung Ching Ho , Choo-Yee Ting

Background Advancements in Large Language Models (LLMs) hold transformative potential in healthcare, however, recent work has raised concern about the tendency of these models to produce outputs that display racial or gender biases.…

Text representation models are prone to exhibit a range of societal biases, reflecting the non-controlled and biased nature of the underlying pretraining data, which consequently leads to severe ethical issues and even bias amplification.…

计算与语言 · 计算机科学 2021-06-08 Soumya Barikeri , Anne Lauscher , Ivan Vulić , Goran Glavaš

The past several years have witnessed a huge surge in the use of social media platforms during mass convergence events such as health emergencies, natural or human-induced disasters. These non-traditional data sources are becoming vital for…

社会与信息网络 · 计算机科学 2023-06-05 Umair Qazi , Muhammad Imran , Ferda Ofli

Text analysis of social media for sentiment, topic analysis, and other analysis depends initially on the selection of keywords and phrases that will be used to create the research corpora. However, keywords that researchers choose may occur…

计算与语言 · 计算机科学 2022-04-21 Philip Feldman , Aaron Dant , James R. Foulds , Shemei Pan

The use of Large Language Models (LLMs) has proven to be a tool that could help in the automatic detection of sexism. Previous studies have shown that these models contain biases that do not accurately reflect reality, especially for…

计算与语言 · 计算机科学 2025-08-26 Judith Tavarez-Rodríguez , Fernando Sánchez-Vega , A. Pastor López-Monroy

Pretrained language models, especially masked language models (MLMs) have seen success across many NLP tasks. However, there is ample evidence that they use the cultural biases that are undoubtedly present in the corpora they are trained…

计算与语言 · 计算机科学 2020-10-02 Nikita Nangia , Clara Vania , Rasika Bhalerao , Samuel R. Bowman

Large language models (LLMs) are known to generate biased responses where the opinions of certain groups and populations are underrepresented. Here, we present a novel approach to achieve controllable generation of specific viewpoints using…

计算与语言 · 计算机科学 2024-04-04 Junyi Li , Ninareh Mehrabi , Charith Peris , Palash Goyal , Kai-Wei Chang , Aram Galstyan , Richard Zemel , Rahul Gupta

Geolocation, the task of identifying an image's location, requires complex reasoning and is crucial for navigation, monitoring, and cultural preservation. However, current methods often produce coarse, imprecise, and non-interpretable…

计算机视觉与模式识别 · 计算机科学 2026-01-07 Zirui Song , Jingpu Yang , Yuan Huang , Jonathan Tonglet , Zeyu Zhang , Tao Cheng , Meng Fang , Iryna Gurevych , Xiuying Chen

This paper presents an overview of a program designed to address the growing need for developing freely available speech resources for under-represented languages. At present we have released 38 datasets for building text-to-speech and…

In the widely used message platform Twitter, about 2% of the tweets contains the geographical location through exact GPS coordinates (latitude and longitude). Knowing the location of a tweet is useful for many data analytics questions. This…

社会与信息网络 · 计算机科学 2015-08-12 Han van der Veen , Djoerd Hiemstra , Tijs van den Broek , Michel Ehrenhard , Ariana Need

A community needs assessment is a tool used by non-profits and government agencies to quantify the strengths and issues of a community, allowing them to allocate their resources better. Such approaches are transitioning towards leveraging…

计算机与社会 · 计算机科学 2024-03-21 Md Towhidul Absar Chowdhury , Naveen Sharma , Ashiqur R. KhudaBukhsh

We perform a large-scale analysis of language diatopic variation using geotagged microblogging datasets. By collecting all Twitter messages written in Spanish over more than two years, we build a corpus from which a carefully selected list…

物理与社会 · 物理学 2014-11-20 Bruno Gonçalves , David Sánchez

Democratizing access to natural language processing (NLP) technology is crucial, especially for underrepresented and extremely low-resource languages. Previous research has focused on developing labeled and unlabeled corpora for these…