中文
相关论文

相关论文: LLM Personas as a Substitute for Field Experiments…

200 篇论文

We investigate the degree to which human plausibility judgments of multiple-choice commonsense benchmark answers are subject to influence by (im)plausibility arguments for or against an answer, in particular, using rationales generated by…

计算与语言 · 计算机科学 2026-02-25 Shramay Palta , Peter Rankel , Sarah Wiegreffe , Rachel Rudinger

Recommender systems are central to online services, enabling users to navigate through massive amounts of content across various domains. However, their evaluation remains challenging due to the disconnect between offline metrics and online…

信息检索 · 计算机科学 2026-04-14 Nicolas Bougie , Gian Maria Marconi , Xiaotong Ye , Narimasa Watanabe

We present a principled approach to provide LLM-based evaluation with a rigorous guarantee of human agreement. We first propose that a reliable evaluation method should not uncritically rely on model preferences for pairwise evaluation, but…

机器学习 · 计算机科学 2024-07-29 Jaehun Jung , Faeze Brahman , Yejin Choi

Large Language Models (LLMs) are increasingly employed for simulating human behaviors across diverse domains. However, our position is that current LLM-based human simulations remain insufficiently reliable, as evidenced by significant…

计算与语言 · 计算机科学 2025-12-02 Qian Wang , Jiaying Wu , Zichen Jiang , Zhenheng Tang , Bingqiao Luo , Nuo Chen , Wei Chen , Bingsheng He

Large language models (LLMs) make it possible to generate synthetic behavioural data at scale, offering an ethical and low-cost alternative to human experiments. Whether such data can faithfully capture psychological differences driven by…

计算与语言 · 计算机科学 2025-11-27 Manuel Pratelli , Marinella Petrocchi

Persona conditioning can be viewed as a behavioral prior for large language models (LLMs) and is often assumed to confer expertise and improve safety in a monotonic manner. However, its effects on high-stakes clinical decision-making remain…

Reliable evaluation is crucial for advancing Automated Program Repair (APR), but prevailing benchmarks rely on execution-based evaluation methods (unit test pass@k), which fail to capture true patch validity. Determining validity can…

软件工程 · 计算机科学 2025-11-17 Sherry Shi , Renyao Wei , Michele Tufano , José Cambronero , Runxiang Cheng , Franjo Ivančić , Pat Rondon

Large language models are increasingly used to represent human opinions, values, or beliefs, and their steerability towards these ideals is an active area of research. Existing work focuses predominantly on aligning marginal response…

计算与语言 · 计算机科学 2026-04-22 Tristan Williams , Franziska Weeber , Sebastian Padó , Alan Akbik

Do horror writers have worse childhoods than other writers? Though biographical details are known about many writers, quantitatively exploring such a qualitative hypothesis requires significant human effort, e.g. to sift through many…

人工智能 · 计算机科学 2024-11-28 Miguel Zabaleta , Joel Lehman

Applications based on large language models (LLMs), such as multi-agent simulations, require population diversity among agents. We identify a pervasive failure mode we term \emph{Persona Collapse}: agents each assigned a distinct profile…

计算与语言 · 计算机科学 2026-04-28 Yunze Xiao , Vivienne J. Zhang , Chenghao Yang , Ningshan Ma , Weihao Xuan , Jen-tse Huang

Recent claims of strong performance by Large Language Models (LLMs) on causal discovery are undermined by a key flaw: many evaluations rely on benchmarks likely included in pretraining corpora. Thus, apparent success suggests that LLM-only…

Evaluating text summarization has been a challenging task in natural language processing (NLP). Automatic metrics which heavily rely on reference summaries are not suitable in many situations, while human evaluation is time-consuming and…

计算与语言 · 计算机科学 2024-07-02 Huyen Nguyen , Haihua Chen , Lavanya Pobbathi , Junhua Ding

Large language models (LLMs), a recent advance in deep learning and machine intelligence, have manifested astonishing capacities, now considered among the most promising for artificial general intelligence. With human-like capabilities,…

人工智能 · 计算机科学 2025-09-19 Zhilun Zhou , Jing Yi Wang , Nicholas Sukiennik , Chen Gao , Fengli Xu , Yong Li , James Evans

Method comparisons are essential to provide recommendations and guidance for applied researchers, who often have to choose from a plethora of available approaches. While many comparisons exist in the literature, these are often not neutral…

统计方法学 · 统计学 2022-12-07 Sarah Friedrich , Tim Friede

Usability evaluation is crucial in human-centered design but can be costly, requiring expert time and user compensation. In this work, we developed a method for synthetic heuristic evaluation using multimodal LLMs' ability to analyze images…

人机交互 · 计算机科学 2025-07-04 Ruican Zhong , David W. McDonald , Gary Hsieh

Large Language Models (LLMs) especially ChatGPT have produced impressive results in various areas, but their potential human-like psychology is still largely unexplored. Existing works study the virtual personalities of LLMs but rarely…

计算与语言 · 计算机科学 2023-10-16 Haocong Rao , Cyril Leung , Chunyan Miao

Personality traits have long been studied as predictors of human behavior. Recent advances in Large Language Models (LLMs) suggest similar patterns may emerge in artificial systems, with advanced LLMs displaying consistent behavioral…

Personalization plays a critical role in numerous language tasks and applications, since users with the same requirements may prefer diverse outputs based on their individual interests. This has led to the development of various…

计算与语言 · 计算机科学 2024-09-19 Jiongnan Liu , Yutao Zhu , Shuting Wang , Xiaochi Wei , Erxue Min , Yu Lu , Shuaiqiang Wang , Dawei Yin , Zhicheng Dou

Evaluating and ranking the capabilities of different LLMs is crucial for understanding their performance and alignment with human preferences. Due to the high cost and time-consuming nature of human evaluations, an automatic LLM bencher…

计算与语言 · 计算机科学 2025-02-12 Mingqi Gao , Yixin Liu , Xinyu Hu , Xiaojun Wan , Jonathan Bragg , Arman Cohan

AIVisor, an agentic retrieval-augmented LLM for student advising, was used to examine how personalization affects system performance across multiple evaluation dimensions. Using twelve authentic advising questions intentionally designed to…

信息检索 · 计算机科学 2026-05-19 Satyajit Movidi , Stephen Russell