中文
相关论文

相关论文: Validating psychometric survey responses

200 篇论文

A central goal of survey research is to collect robust and reliable data from respondents. However, despite researchers' best efforts in designing questionnaires, respondents may experience difficulty understanding questions' intent and…

人机交互 · 计算机科学 2020-11-16 Amanda Fernández-Fontelo , Pascal J. Kieslich , Felix Henninger , Frauke Kreuter , Sonja Greven

As psychometric surveys are increasingly used to assess the traits of large language models (LLMs), the need for scalable survey item generation suited for LLMs has also grown. A critical challenge here is ensuring the construct validity of…

计算与语言 · 计算机科学 2026-05-26 Sungjib Lim , Woojung Song , Eun-Ju Lee , Yohan Jo

Large-scale surveys are essential tools for informing social science research and policy, but running surveys is costly and time-intensive. If we could accurately simulate group-level survey results, this would therefore be very valuable to…

计算与语言 · 计算机科学 2025-02-20 Yong Cao , Haijiang Liu , Arnav Arora , Isabelle Augenstein , Paul Röttger , Daniel Hershcovich

Assessing the validity of user simulators when used for the evaluation of information retrieval systems remains an open question, constraining their effective use and the reliability of simulation-based results. To address this issue, we…

信息检索 · 计算机科学 2026-01-19 Andreas Konstantin Kruff , Nolwenn Bernard , Philipp Schaer

ChatGPT and other large language models (LLMs) have proven useful in crowdsourcing tasks, where they can effectively annotate machine learning training data. However, this means that they also have the potential for misuse, specifically to…

Survey research is a fundamental empirical method in software engineering, enabling the systematic collection of data on professional practices, perceptions, and experiences. However, recent advances in large language models (LLMs) have…

Large language models (LLMs) are increasingly used to simulate human opinions and survey responses, but their ability to reproduce population responses across cultures remains limited. Existing persona-based prompting methods typically rely…

计算与语言 · 计算机科学 2026-05-18 Axel Abels , Elias Fernandez Domingos , Apurva Shah , Tom Lenaerts

A growing literature uses large language models (LLMs) as synthetic participants to generate cost-effective and nearly instantaneous responses in social science experiments. However, there is limited guidance on when such simulations…

人工智能 · 计算机科学 2026-02-18 Jessica Hullman , David Broska , Huaman Sun , Aaron Shaw

The validity of online behavioral research relies on study participants being human rather than machine. In the past, it was possible to detect machines by posing simple challenges that were easily solved by humans but not by machines.…

计算与语言 · 计算机科学 2026-04-02 Simon Schug , Brenden M. Lake

Large language models (LLMs) increasingly reach real-world applications, necessitating a better understanding of their behaviour. Their size and complexity complicate traditional assessment methods, causing the emergence of alternative…

人工智能 · 计算机科学 2025-05-13 Sanne Peereboom , Inga Schwabe , Bennett Kleinberg

Difficulty spillover and suboptimal help-seeking challenge the sequential, knowledge-intensive nature of digital tasks. In online surveys, tough questions can drain mental energy and hurt performance on later questions, while users often…

人机交互 · 计算机科学 2026-02-04 Ailin Liu , Yesmine Karoui , Fiona Draxler , Frauke Kreuter , Francesco Chiossi

Large language models (LLMs) are increasingly used to simulate survey responses, but synthetic data can be misaligned with the human population, leading to unreliable inference. We develop a general framework that converts LLM-simulated…

统计方法学 · 统计学 2026-05-21 Chengpiao Huang , Yuhang Wu , Kaizheng Wang

Digital personas powered by Large Language Models (LLMs) are increasingly proposed as substitutes for human survey respondents, yet it remains unclear when they can reliably approximate human survey findings. We answer this question using…

计算与语言 · 计算机科学 2026-05-12 Mumin Jia , Yilin Chen , Divya Sharma , Jairo Diaz-Rodriguez

Using Large Language Models (LLMs) to simulate user opinions has received growing attention. Yet LLMs, especially trained with reinforcement learning from human feedback (RLHF), are known to exhibit biases toward dominant viewpoints,…

计算与语言 · 计算机科学 2025-12-09 Ziyun Yu , Yiru Zhou , Chen Zhao , Hongyi Wen

In recent years, the use of machine learning classifiers is of great value in solving a variety of problems in text classification. Sentiment mining is a kind of text classification in which, messages are classified according to sentiment…

机器学习 · 计算机科学 2014-02-18 Vinodhini G Chandrasekaran RM

In this paper we describe the usefulness of statistical validation techniques for human factors survey research. We need to investigate a diversity of validity aspects when creating metrics in human factors research, and we argue that the…

软件工程 · 计算机科学 2019-04-05 Lucas Gren , Alfredo Goldman

Large language models (LLMs) are increasingly used as proxies for human judgment in computational social science, yet their ability to reproduce patterns of susceptibility to misinformation remains unclear. We test whether LLM-simulated…

社会与信息网络 · 计算机科学 2026-04-13 Eun Cheol Choi , Lindsay E. Young , Emilio Ferrara

Language models (LMs) are increasingly used to simulate human-like responses in scenarios where accurately mimicking a population's behavior can guide decision-making, such as in developing educational materials and designing public…

计算与语言 · 计算机科学 2024-07-23 Joy He-Yueya , Wanjing Anya Ma , Kanishk Gandhi , Benjamin W. Domingue , Emma Brunskill , Noah D. Goodman

Real-world applications of machine learning models are often subject to legal or policy-based regulations. Some of these regulations require ensuring the validity of the model, i.e., the approximation error being smaller than a threshold. A…

机器学习 · 统计学 2024-06-18 Sven Lämmle , Can Bogoclu , Robert Voßhall , Anselm Haselhoff , Dirk Roos

Psychometric tests are increasingly used to assess psychological constructs in large language models (LLMs). However, it remains unclear whether these tests -- originally developed for humans -- yield meaningful results when applied to…

计算与语言 · 计算机科学 2026-01-28 Jana Jung , Marlene Lutz , Indira Sen , Markus Strohmaier
‹ 上一页 1 2 3 10 下一页 ›