中文
相关论文

相关论文: LLM Personas as a Substitute for Field Experiments…

200 篇论文

Many evaluations of large language models (LLMs) in text annotation focus primarily on the correctness of the output, typically comparing model-generated labels to human-annotated ``ground truth'' using standard performance metrics. In…

信息检索 · 计算机科学 2025-10-30 Jiaman He , Zikang Leng , Dana McKay , Damiano Spina , Johanne R. Trippas

In NLG meta-evaluation, evaluation metrics are typically assessed based on their consistency with humans. However, we identify some limitations in traditional NLG meta-evaluation approaches, such as issues in handling human ratings and…

计算与语言 · 计算机科学 2025-08-18 Xinyu Hu , Mingqi Gao , Li Lin , Zhenghan Yu , Xiaojun Wan

As large language models (LLMs) are increasingly used in human-centered tasks, assessing their psychological traits is crucial for understanding their social impact and ensuring trustworthy AI alignment. While existing reviews have covered…

Deploying large language models (LLMs) with agency in real-world applications raises critical questions about how these models will behave. In particular, how will their decisions align with humans when faced with moral dilemmas? This study…

计算机与社会 · 计算机科学 2025-04-16 Jiseon Kim , Jea Kwon , Luiz Felipe Vecchietti , Alice Oh , Meeyoung Cha

Large language models (LLMs) are increasingly used in the social sciences to simulate human behavior, based on the assumption that they can generate realistic, human-like text. Yet this assumption remains largely untested. Existing…

计算与语言 · 计算机科学 2025-11-26 Nicolò Pagan , Petter Törnberg , Christopher A. Bail , Anikó Hannák , Christopher Barrie

Large language models (LLMs) have advanced conversational AI assistants. However, systematically evaluating how well these assistants apply personalization--adapting to individual user preferences while completing tasks--remains…

计算与语言 · 计算机科学 2025-06-12 Zheng Zhao , Clara Vania , Subhradeep Kayal , Naila Khan , Shay B. Cohen , Emine Yilmaz

The humanlike responses of large language models (LLMs) have prompted social scientists to investigate whether LLMs can be used to simulate human participants in experiments, opinion polls and surveys. Of central interest in this line of…

计算与语言 · 计算机科学 2024-05-14 Nikolay B Petrov , Gregory Serapio-García , Jason Rentfrow

Current role-play studies often rely on unvalidated LLM-as-a-judge paradigms, which may fail to reflect how humans perceive role fidelity. A key prerequisite for human-aligned evaluation is role identification, the ability to recognize who…

计算与语言 · 计算机科学 2025-08-15 Lingfeng Zhou , Jialing Zhang , Jin Gao , Mohan Jiang , Dequan Wang

The last couple of years have witnessed emerging research that appropriates Theory-of-Mind (ToM) tasks designed for humans to benchmark LLM's ToM capabilities as an indication of LLM's social intelligence. However, this approach has a…

人机交互 · 计算机科学 2025-04-16 Qiaosi Wang , Xuhui Zhou , Maarten Sap , Jodi Forlizzi , Hong Shen

Large language models (LLMs) offer emerging opportunities for psychological and behavioral research, but methodological guidance is lacking. This article provides a framework for using LLMs as psychological simulators across two primary…

计算机与社会 · 计算机科学 2026-04-07 Zhicheng Lin

As language models achieve increasingly human-like capabilities in conversational text generation, a critical question emerges: to what extent can these systems simulate the characteristics of specific individuals? To evaluate this, we…

计算与语言 · 计算机科学 2025-06-04 Quan Shi , Carlos E. Jimenez , Stephen Dong , Brian Seo , Caden Yao , Adam Kelch , Karthik Narasimhan

As evaluation designs of large language models may shape our trajectory toward artificial general intelligence, comprehensive and forward-looking assessment is essential. Existing benchmarks primarily assess static knowledge, while…

计算与语言 · 计算机科学 2025-08-07 Jiayin Wang , Zhiquang Guo , Weizhi Ma , Min Zhang

Relevance evaluation plays a crucial role in personalized search systems to ensure that search results align with a user's queries and intent. While human annotation is the traditional method for relevance evaluation, its high cost and long…

信息检索 · 计算机科学 2025-11-12 Han Wang , Alex Whitworth , Pak Ming Cheung , Zhenjie Zhang , Krishna Kamath

While rapid advances in large language models (LLMs) are reshaping data-driven intelligent education, accurately simulating students remains an important but challenging bottleneck for scalable educational data collection, evaluation, and…

计算机与社会 · 计算机科学 2025-12-05 Haoxuan Li , Jifan Yu , Xin Cong , Yang Dang , Daniel Zhang-li , Lu Mi , Yisi Zhan , Huiqin Liu , Zhiyuan Liu

Aligning LLM-based judges with human preferences is a significant challenge, as they are difficult to calibrate and often suffer from rubric sensitivity, bias, and instability. Overcoming this challenge advances key applications, such as…

LLM use in annotation is becoming widespread, and given LLMs' overall promising performance and speed, simply "reviewing" LLM annotations in interpretive tasks can be tempting. In subjective annotation tasks with multiple plausible answers,…

计算机与社会 · 计算机科学 2025-07-22 Hope Schroeder , Deb Roy , Jad Kabbara

Personalized AI agents are becoming central to modern information retrieval, yet most evaluation methodologies remain static, relying on fixed benchmarks and one-off metrics that fail to reflect how users' needs evolve over time. These…

信息检索 · 计算机科学 2025-10-07 Kirandeep Kaur , Preetam Prabhu Srikar Dammu , Hideo Joho , Chirag Shah

Large language models (LLMs) have demonstrated unprecedented emergent capabilities, including content generation, translation, and simulation of human behavior. Field experiments, on the other hand, are widely employed in social studies to…

计算机与社会 · 计算机科学 2025-05-22 Yaoyu Chen , Yuheng Hu , Yingda Lu

Large language models (LLMs) have demonstrated significant potential in developing Role-Playing Agents (RPAs). However, current research primarily evaluates RPAs using famous fictional characters, allowing models to rely on memory…

计算与语言 · 计算机科学 2026-03-05 Ji-Lun Peng , Yun-Nung Chen

Artificial General Intelligence falls short when communicating role specific nuances to other systems. This is more pronounced when building autonomous LLM agents capable and designed to communicate with each other for real world problem…

机器学习 · 计算机科学 2024-03-19 Rabimba Karanjai , Weidong Shi