English
Related papers

Related papers: SocioEval: A Template-Based Framework for Evaluati…

200 papers

Large language models (LLMs) are now widely deployed in user-facing applications, reaching hundreds of millions worldwide. As they become integrated into everyday tasks, growing reliance on their outputs raises significant concerns. In…

Computers and Society · Computer Science 2025-10-16 Robin Staab , Jasper Dekoninck , Maximilian Baader , Martin Vechev

Recently, there has been growing interest in extending the context length of large language models (LLMs), aiming to effectively process long inputs of one turn or conversations with more extensive histories. While proprietary models such…

Computation and Language · Computer Science 2023-10-05 Chenxin An , Shansan Gong , Ming Zhong , Xingjian Zhao , Mukai Li , Jun Zhang , Lingpeng Kong , Xipeng Qiu

As large language models (LLMs) are employed worldwide, existing evaluation paradigms for their multilingual capabilities primarily focus on factual task performance, neglecting the ability to judge content's deep-level values across…

Computation and Language · Computer Science 2026-05-12 Yukun Chen , Xinyu Zhang , Boyi Deng , Jialong Tang , Yu Wan , Fei Huang , Yuxi Zhou , Baosong Yang , Yiming Li

In the rapidly evolving field of artificial intelligence, large language models (LLMs) have emerged as powerful tools for a myriad of applications, from natural language processing to decision-making support systems. However, as these…

Computation and Language · Computer Science 2025-07-08 Jianchao Ji , Yutong Chen , Mingyu Jin , Wujiang Xu , Wenyue Hua , Yongfeng Zhang

Large Language Models (LLMs) are increasingly deployed in resume screening pipelines. Although explicit PII (e.g., names) is commonly redacted, resumes typically retain subtle sociocultural markers (languages, co-curricular activities,…

Computers and Society · Computer Science 2026-05-06 Bryan Chen Zhengyu Tan , Shaun Khoo , Bich Ngoc Doan , Zhengyuan Liu , Nancy F. Chen , Roy Ka-Wei Lee

Recent advances in Large Language Models (LLMs) have enabled human-like responses across various tasks, raising questions about their ethical decision-making capabilities and potential biases. This study systematically evaluates how nine…

Computers and Society · Computer Science 2025-11-03 Wentao Xu , Yile Yan , Yuqi Zhu

Evaluating alignment in language models requires testing how they behave under realistic pressure, not just what they claim they would do. While alignment failures increasingly cause real-world harm, comprehensive evaluation frameworks with…

Artificial Intelligence · Computer Science 2026-02-25 Nora Petrova , John Burden

For Large Language Models (LLMs), a disconnect persists between benchmark performance and real-world utility. Current evaluation frameworks remain fragmented, prioritizing technical metrics while neglecting holistic assessment for…

Artificial Intelligence · Computer Science 2025-11-19 Jun Wang , Ninglun Gu , Kailai Zhang , Zijiao Zhang , Yelun Bao , Jin Yang , Xu Yin , Liwei Liu , Yihuan Liu , Pengyong Li , Gary G. Yen , Junchi Yan

As Vision-Language Models (VLMs) become integral to educational decision-making, ensuring their fairness is paramount. However, current text-centric evaluations neglect the visual modality, leaving an unregulated channel for latent social…

Artificial Intelligence · Computer Science 2026-04-15 Ruijia Li , Mingzi Zhang , Zengyi Yu , Yuang Wei , Bo Jiang

Social categories and stereotypes are embedded in language and can introduce data bias into Large Language Models (LLMs). Despite safeguards, these biases often persist in model behavior, potentially leading to representational harm in…

Computation and Language · Computer Science 2025-02-27 Rebekka Görge , Michael Mock , Héctor Allende-Cid

Recent advancements in Large Language Models (LLMs) have positioned them as powerful tools for clinical decision-making, with rapidly expanding applications in healthcare. However, concerns about bias remain a significant challenge in the…

Artificial Intelligence · Computer Science 2024-10-23 Kenza Benkirane , Jackie Kay , Maria Perez-Ortiz

Large language models (LLMs) are increasingly trained from AI constitutions and model specifications that establish behavioral guidelines and ethical principles. However, these specifications face critical challenges, including internal…

Computation and Language · Computer Science 2025-10-24 Jifan Zhang , Henry Sleight , Andi Peng , John Schulman , Esin Durmus

The advances of large foundation models necessitate wide-coverage, low-cost, and zero-contamination benchmarks. Despite continuous exploration of language model evaluations, comprehensive studies on the evaluation of Large Multi-modal…

Computation and Language · Computer Science 2025-09-19 Kaichen Zhang , Bo Li , Peiyuan Zhang , Fanyi Pu , Joshua Adrian Cahyono , Kairui Hu , Shuai Liu , Yuanhan Zhang , Jingkang Yang , Chunyuan Li , Ziwei Liu

Large language models (LLMs) are widely applied across diverse domains, raising concerns about their limitations and potential risks. In this study, we investigate two types of bias that LLMs may display: stereotype bias and deviation bias.…

Computation and Language · Computer Science 2026-05-20 Daniel Wang , Eli Brignac , Minjia Mao , Xiao Fang

As modern Large Language Models (LLMs) shatter many state-of-the-art benchmarks in a variety of domains, this paper investigates their behavior in the domains of ethics and fairness, focusing on protected group bias. We conduct a two-part…

Computers and Society · Computer Science 2024-03-25 Hadas Kotek , David Q. Sun , Zidi Xiu , Margit Bowler , Christopher Klein

Artificial Intelligence (AI) is increasingly used in hiring, with large language models (LLMs) having the potential to influence or even make hiring decisions. However, this raises pressing concerns about bias, fairness, and trust,…

Computers and Society · Computer Science 2025-08-26 Pooja S. B. Rao , Laxminarayen Nagarajan Venkatesan , Mauro Cherubini , Dinesh Babu Jayagopi

This paper introduces a comprehensive benchmark for evaluating how Large Language Models (LLMs) respond to linguistic shibboleths: subtle linguistic markers that can inadvertently reveal demographic attributes such as gender, social class,…

Computation and Language · Computer Science 2025-08-08 Julia Kharchenko , Tanya Roosta , Aman Chadha , Chirag Shah

The use of Large Language Models (LLMs) has proven to be a tool that could help in the automatic detection of sexism. Previous studies have shown that these models contain biases that do not accurately reflect reality, especially for…

Computation and Language · Computer Science 2025-08-26 Judith Tavarez-Rodríguez , Fernando Sánchez-Vega , A. Pastor López-Monroy

Large language models (LLMs) are now deployed worldwide, inspiring a surge of benchmarks that measure their multilingual and multicultural abilities. However, these benchmarks prioritize generic language understanding or superficial…

Despite the recent strides in large language models, studies have underscored the existence of social biases within these systems. In this paper, we delve into the validation and comparison of the ethical biases of LLMs concerning globally…

Computation and Language · Computer Science 2025-07-03 Seunguk Yu , Juhwan Choi , Youngbin Kim