English
Related papers

Related papers: Robust Evaluation Measures for Evaluating Social B…

200 papers

There has been significant prior work using templates to study bias against demographic attributes in MLMs. However, these have limitations: they overlook random variability of templates and target concepts analyzed, assume equality amongst…

Computation and Language · Computer Science 2025-08-25 Ingroj Shrestha , Louis Tay , Padmini Srinivasan

While generative multilingual models are rapidly being deployed, their safety and fairness evaluations are largely limited to resources collected in English. This is especially problematic for evaluations targeting inherently socio-cultural…

Computation and Language · Computer Science 2024-03-12 Mukul Bhutani , Kevin Robinson , Vinodkumar Prabhakaran , Shachi Dave , Sunipa Dev

Language models (LMs) are increasingly used as simulacra for people, yet their ability to match the distribution of views of a specific demographic group and be \textit{distributionally aligned} remains uncertain. This notion of…

Computation and Language · Computer Science 2024-11-11 Nicole Meister , Carlos Guestrin , Tatsunori Hashimoto

The relationship between communicated language and intended meaning is often probabilistic and sensitive to context. Numerous strategies attempt to estimate such a mapping, often leveraging recursive Bayesian models of communication. In…

Computation and Language · Computer Science 2023-05-03 Benjamin Lipkin , Lionel Wong , Gabriel Grand , Joshua B Tenenbaum

Large language models (LLMs) have revolutionised many fields, with LLM-as-a-service (LLMSaaS) offering accessible, general-purpose solutions without costly task-specific training. In contrast to the widely studied prompt engineering for…

Computation and Language · Computer Science 2026-03-17 Md Rakibul Hasan , Yue Yao , Md Zakir Hossain , Aneesh Krishna , Imre Rudas , Shafin Rahman , Tom Gedeon

Evaluations are critical for understanding the capabilities of large language models (LLMs). Fundamentally, evaluations are experiments; but the literature on evaluations has largely ignored the literature from other sciences on experiment…

Applications · Statistics 2024-11-04 Evan Miller

The Kullback-Leibler (KL) divergence is frequently used in data science. For discrete distributions on large state spaces, approximations of probability vectors may result in a few small negative entries, rendering the KL divergence…

As Large Language Models (LLMs) increasingly appear in social science research (e.g., economics and marketing), it becomes crucial to assess how well these models replicate human behavior. In this work, using hypothesis testing, we present…

Computers and Society · Computer Science 2025-06-19 Harbin Hong , Sebastian Caldas , Liu Leqi

Multi-modal Large Language Models (MLLMs) have dramatically advanced the research field and delivered powerful vision-language understanding capabilities. However, these models often inherit deep-rooted social biases from their training…

Computation and Language · Computer Science 2025-08-21 Harry Cheng , Yangyang Guo , Qingpei Guo , Ming Yang , Tian Gan , Weili Guan , Liqiang Nie

Recently, language models have demonstrated strong performance on various natural language understanding tasks. Language models trained on large human-generated corpus encode not only a significant amount of human knowledge, but also the…

Computation and Language · Computer Science 2023-04-10 Damin Zhang , Julia Rayz , Romila Pradhan

Multimodal Large Language Models (MLLMs) are increasingly applied in Personalized Image Aesthetic Assessment (PIAA) as a scalable alternative to expert evaluations. However, their predictions may reflect subtle biases influenced by…

Computation and Language · Computer Science 2025-09-16 Kun Li , Lai-Man Po , Hongzheng Yang , Xuyuan Xu , Kangcheng Liu , Yuzhi Zhao

Consumer research costs companies billions annually yet suffers from panel biases and limited scale. Large language models (LLMs) offer an alternative by simulating synthetic consumers, but produce unrealistic response distributions when…

Artificial Intelligence · Computer Science 2025-10-28 Benjamin F. Maier , Ulf Aslak , Luca Fiaschi , Nina Rismal , Kemble Fletcher , Christian C. Luhmann , Robbie Dow , Kli Pappas , Thomas V. Wiecki

Large language models (LLMs) are known to inherit and even amplify societal biases present in their pre-training corpora, threatening fairness and social trust. To address this issue, recent work has explored ``editing'' LLM parameters to…

Computation and Language · Computer Science 2025-12-03 Daiki Shirafuji , Tatsuhiko Saito , Yasutomo Kimura

Stereotype detection is a challenging and subjective task, as certain statements, such as "Black people like to play basketball," may not appear overtly toxic but still reinforce racial stereotypes. With the increasing prevalence of large…

Computation and Language · Computer Science 2024-11-19 Zekun Wu , Sahan Bulathwela , Maria Perez-Ortiz , Adriano Soares Koshiyama

Marginal maximum likelihood estimation (MMLE) in item response theory (IRT) is highly sensitive to aberrant responses, such as careless answering and random guessing, which can reduce estimation accuracy. To address this issue, this study…

Methodology · Statistics 2025-02-18 Yuki Itaya , Kenichi Hayashi

Numerous types of social biases have been identified in pre-trained language models (PLMs), and various intrinsic bias evaluation measures have been proposed for quantifying those social biases. Prior works have relied on human annotated…

Computation and Language · Computer Science 2023-01-31 Masahiro Kaneko , Danushka Bollegala , Naoaki Okazaki

In this position paper, we argue that human evaluation of generative large language models (LLMs) should be a multidisciplinary undertaking that draws upon insights from disciplines such as user experience research and human behavioral…

Computation and Language · Computer Science 2024-10-08 Aparna Elangovan , Ling Liu , Lei Xu , Sravan Bodapati , Dan Roth

The global increase in mental illness requires innovative detection methods for early intervention. Social media provides a valuable platform to identify mental illness through user-generated content. This systematic review examines machine…

Machine Learning · Computer Science 2025-02-18 Yuchen Cao , Jianglai Dai , Zhongyan Wang , Yeyubei Zhang , Xiaorui Shen , Yunchong Liu , Yexin Tian

Large-scale web-scraped text corpora used to train general-purpose AI models often contain harmful demographic-targeted social biases, creating a regulatory need for data auditing and developing scalable bias-detection methods. Although…

Computation and Language · Computer Science 2026-04-10 Ayan Majumdar , Feihao Chen , Jinghui Li , Xiaozhen Wang

Multilingual Large Language Models (LLMs) exhibit remarkable cross-lingual abilities, yet often exhibit a systematic bias toward the representations from other languages, resulting in semantic interference when generating content in…

Computation and Language · Computer Science 2026-01-21 Ilia Badanin , Daniil Dzenhaliou , Imanol Schlag
‹ Prev 1 4 5 6 7 8 10 Next ›