中文
相关论文

相关论文: Validating psychometric survey responses

200 篇论文

Confidence in LLMs is a useful indicator of model uncertainty and answer reliability. Existing work mainly focused on single-turn scenarios, while research on confidence in complex multi-turn interactions is limited. In this paper, we…

计算与语言 · 计算机科学 2025-10-29 Litu Ou , Kuan Li , Huifeng Yin , Liwen Zhang , Zhongwang Zhang , Xixi Wu , Rui Ye , Zile Qiao , Pengjun Xie , Jingren Zhou , Yong Jiang

Accurate and computationally efficient means for classifying human activities have been the subject of extensive research efforts. Most current research focuses on extracting complex features to achieve high classification accuracy. We…

人工智能 · 计算机科学 2015-12-22 Skyler Seto , Wenyu Zhang , Yichen Zhou

We consider the problem of classifying documents not by topic, but by overall sentiment, e.g., determining whether a review is positive or negative. Using movie reviews as data, we find that standard machine learning techniques definitively…

计算与语言 · 计算机科学 2007-05-23 Bo Pang , Lillian Lee , Shivakumar Vaithyanathan

Social surveys have been widely used as a method of obtaining public opinion. Sometimes it is more ideal to collect opinions by presenting questions in free-response formats than in multiple-choice formats. Despite their advantages,…

社会与信息网络 · 计算机科学 2020-07-09 Tatsuro Kawamoto , Takaaki Aoki

System-oriented IR evaluations are limited to rather abstract understandings of real user behavior. As a solution, simulating user interactions provides a cost-efficient way to support system-oriented experiments with more realistic…

信息检索 · 计算机科学 2022-03-25 Timo Breuer , Norbert Fuhr , Philipp Schaer

In many areas of data mining, data is collected from humans beings. In this contribution, we ask the question of how people actually respond to ordinal scales. The main problem observed is that users tend to be volatile in their choices,…

人机交互 · 计算机科学 2017-03-01 Kevin Jasberg , Sergej Sizov

Automated fault localization is an important issue in model validation and verification. It helps the end users in analyzing the origin of failure. In this work, we show the early experiments with probabilistic analysis approaches in fault…

软件工程 · 计算机科学 2016-11-21 Ning Ge , Marc Pantel , Xavier Crégut

This study provides evidence that personality can be reliably predicted from activity data collected through mobile phone sensors. Employing a set of well informed indicators calculable from accelerometer records and movement patterns, we…

信号处理 · 电气工程与系统科学 2024-01-23 Wun Yung Shaney Sze , Maryglen Pearl Herrero , Roger Garriga

Smartphones have ubiquitously integrated into our home and work environments, however, users normally rely on explicit but inefficient identification processes in a controlled environment. Therefore, when a device is stolen, a thief can…

密码学与安全 · 计算机科学 2020-09-25 Muhammad Ahmad , Ali Kashif Bashir , Adil Mehmood Khan , Manuel Mazzara , Salvatore Distefano , Shahzad Sarfraz

Surveys have recently gained popularity as a tool to study large language models. By comparing survey responses of models to those of human reference populations, researchers aim to infer the demographics, political opinions, or values best…

计算与语言 · 计算机科学 2024-12-10 Ricardo Dominguez-Olmedo , Moritz Hardt , Celestine Mendler-Dünner

Motivated by the remarkable progress of large language models (LLMs) in objective tasks like mathematics and coding, there is growing interest in their potential to simulate human behavior--a capability with profound implications for…

计算与语言 · 计算机科学 2026-01-23 Yuxuan Lei , Tianfu Wang , Jianxun Lian , Zhengyu Hu , Defu Lian , Xing Xie

The recent development and wider accessibility of LLMs have spurred discussions about how they can be used in survey research, including classifying open-ended survey responses. Due to their linguistic capacities, it is possible that LLMs…

计算与语言 · 计算机科学 2025-07-04 Leah von der Heyde , Anna-Carolina Haensch , Bernd Weiß , Jessica Daikeler

The ability of Large Language Models (LLMs) to mimic human behavior triggered a plethora of computational social science research, assuming that empirical studies of humans can be conducted with AI agents instead. Since there have been…

计算与语言 · 计算机科学 2025-10-21 Simon Münker , Nils Schwager , Achim Rettinger

Automatic evaluation of various text quality criteria produced by data-driven intelligent methods is very common and useful because it is cheap, fast, and usually yields repeatable results. In this paper, we present an attempt to automate…

计算与语言 · 计算机科学 2020-06-08 Erion Çano , Ondřej Bojar

Upsurging abnormal activities in crowded locations such as airports, train stations, bus stops, shopping malls, etc., urges the necessity for an intelligent surveillance system. An intelligent surveillance system can differentiate between…

计算机视觉与模式识别 · 计算机科学 2024-09-04 Shahriar Jahan , Roknuzzaman , Md Robiul Islam

Auditing the use of data in training machine-learning (ML) models is an increasingly pressing challenge, as myriad ML practitioners routinely leverage the effort of content creators to train models without their permission. In this paper,…

密码学与安全 · 计算机科学 2025-01-28 Zonghao Huang , Neil Zhenqiang Gong , Michael K. Reiter

This paper tackles the challenging task of evaluating socially situated conversational robots and presents a novel objective evaluation approach that relies on multimodal user behaviors. In this study, our main focus is on assessing the…

计算与语言 · 计算机科学 2023-09-26 Koji Inoue , Divesh Lala , Keiko Ochi , Tatsuya Kawahara , Gabriel Skantze

As we consider entrusting Large Language Models (LLMs) with key societal and decision-making roles, measuring their alignment with human cognition becomes critical. This requires methods that can assess how these systems represent…

人工智能 · 计算机科学 2025-10-03 Mattson Ogg , Ritwik Bose , Jamie Scharf , Christopher Ratto , Michael Wolmetz

Large language models (LLM) agents may offer tools to predict human responses to surveys. A common technique for defining these agents uses only demographics, for example country, age, gender, employment status, income, education and…

计算机与社会 · 计算机科学 2026-05-19 Rubén Garzón , Pauline Baron , Vincent Grari , Jonne Kamphorst , Michael Bernstein , Marcin Detyniecki

Consumer research costs companies billions annually yet suffers from panel biases and limited scale. Large language models (LLMs) offer an alternative by simulating synthetic consumers, but produce unrealistic response distributions when…