English
Related papers

Related papers: Quantifying the Utility of User Simulators for Bui…

200 papers

As user simulators are increasingly used for interactive training and evaluation of AI assistants, it is essential that they represent the diverse behaviors of real users. While existing works train user simulators to generate human-like…

Computation and Language · Computer Science 2026-05-11 Shuhaib Mehri , Philippe Laban , Sumuk Shashidhar , Marwa Abdulhai , Sergey Levine , Michel Galley , Dilek Hakkani-Tür

User simulation is a promising approach for automatically training and evaluating conversational information access agents, enabling the generation of synthetic dialogues and facilitating reproducible experiments at scale. However, the…

Information Retrieval · Computer Science 2024-06-28 Nolwenn Bernard , Krisztian Balog

A long-standing challenge in developing accurate recommendation models is simulating user behavior, mainly due to the complex and stochastic nature of user interactions. Towards this, one promising line of work has been the use of Large…

Information Retrieval · Computer Science 2025-09-15 Himanshu Thakur , Eshani Agrawal , Smruthi Mukund

Agent skills, which are reusable, domain-specific knowledge artifacts, have become a popular mechanism for extending LLM-based agents, yet formally benchmarking skill usage performance remains scarce. Existing skill benchmarking efforts…

Computation and Language · Computer Science 2026-04-07 Yujian Liu , Jiabao Ji , Li An , Tommi Jaakkola , Yang Zhang , Shiyu Chang

Truthfulness (adherence to factual accuracy) and utility (satisfying human needs and instructions) are both fundamental aspects of Large Language Models, yet these goals often conflict (e.g., sell a car with known flaws), which makes it…

Artificial Intelligence · Computer Science 2025-04-29 Zhe Su , Xuhui Zhou , Sanketh Rangreji , Anubha Kabra , Julia Mendelsohn , Faeze Brahman , Maarten Sap

Evaluation of large language model (LLM) outputs requires users to make critical judgments about the best outputs across various configurations. This process is costly and takes time given the large amounts of data. LLMs are increasingly…

Modeling subrational agents, such as humans or economic households, is inherently challenging due to the difficulty in calibrating reinforcement learning models or collecting data that involves human subjects. Existing work highlights the…

Artificial Intelligence · Computer Science 2024-02-15 Andrea Coletta , Kshama Dwarakanath , Penghang Liu , Svitlana Vyetrenko , Tucker Balch

Simulation powered by Large Language Models (LLMs) has become a promising method for exploring complex human social behaviors. However, the application of LLMs in simulations presents significant challenges, particularly regarding their…

Computers and Society · Computer Science 2025-02-26 Qian Wang , Zhenheng Tang , Bingsheng He

Large language models (LLMs) are increasingly used to simulate human behavior, but their ability to simulate $individual$ privacy decisions is not well understood. In this paper, we address the problem of evaluating whether a core set of…

Cryptography and Security · Computer Science 2026-05-13 James Flemings , Murali Annavaram

Reward models (RMs) are crucial for the training of large language models (LLMs), yet they typically rely on large-scale human-annotated preference pairs. With the widespread deployment of LLMs, in-the-wild interactions have emerged as a…

Computation and Language · Computer Science 2026-02-10 Hao Peng , Yunjia Qi , Xiaozhi Wang , Zijun Yao , Lei Hou , Juanzi Li

To some, the advent of artificial intelligence (AI) promises better decision-making and increased military effectiveness while reducing the influence of human error and emotions. However, there is still debate about how AI systems,…

Computers and Society · Computer Science 2024-10-04 Max Lamparth , Anthony Corso , Jacob Ganz , Oriana Skylar Mastro , Jacquelyn Schneider , Harold Trinkunas

Training mental health clinicians to conduct standardized clinical assessments is challenging due to a lack of scalable, realistic practice opportunities, which can impact data quality in clinical trials. To address this gap, we introduce a…

Human-Computer Interaction · Computer Science 2025-12-30 Veronica Bossio Botero , Vijay Yadav , Jacob Ouyang , Anzar Abbas , Michelle Worthington

Recent research shows that LLM Agents can generate ``believable'' human behaviors via prompt-only methods, and such agents have been increasingly adopted in downstream applications. However, existing evaluation of these agents only focuses…

Computation and Language · Computer Science 2026-04-30 Yuxuan Lu , Jing Huang , Yan Han , Bingsheng Yao , Sisong Bei , Jiri Gesi , Yaochen Xie , Yisi Sang , Zheshen , Wang , Qi He , Dakuo Wang

User simulators are essential to conversational AI, enabling scalable agent development and evaluation through simulated interactions. While current Large Language Models (LLMs) have advanced user simulation capabilities, we reveal that…

Computation and Language · Computer Science 2026-03-10 Shuhaib Mehri , Xiaocheng Yang , Takyoung Kim , Gokhan Tur , Shikib Mehri , Dilek Hakkani-Tür

Personalisation is a standard feature of conversational AI systems used by millions; yet, the efficacy of personalisation methods is often evaluated in academic research using simulated users rather than real people. This raises questions…

Computation and Language · Computer Science 2026-05-14 Hannah Rose Kirk , Liu Leqi , Fanzhi Zeng , Henry Davidson , Bertie Vidgen , Christopher Summerfield , Scott A. Hale

While rapid advances in large language models (LLMs) are reshaping data-driven intelligent education, accurately simulating students remains an important but challenging bottleneck for scalable educational data collection, evaluation, and…

Computers and Society · Computer Science 2025-12-05 Haoxuan Li , Jifan Yu , Xin Cong , Yang Dang , Daniel Zhang-li , Lu Mi , Yisi Zhan , Huiqin Liu , Zhiyuan Liu

Quality-sensitive applications of machine learning (ML) require quality assurance (QA) by humans before the predictions of an ML model can be deployed. QA for ML (QA4ML) interfaces require users to view a large amount of data and perform…

Human-Computer Interaction · Computer Science 2023-09-01 Yu Zhang , Martijn Tennekes , Tim de Jong , Lyana Curier , Bob Coecke , Min Chen

Large Language Models are typically trained with next-turn rewards, limiting their ability to optimize for long-term interaction. As a result, they often respond passively to ambiguous or open-ended user requests, failing to help users…

Artificial Intelligence · Computer Science 2025-07-31 Shirley Wu , Michel Galley , Baolin Peng , Hao Cheng , Gavin Li , Yao Dou , Weixin Cai , James Zou , Jure Leskovec , Jianfeng Gao

Scalable AI agents training relies on interactive environments that faithfully simulate the consequences of agent actions. Manually crafted environments are expensive to build, brittle to extend, and fundamentally limited in diversity. A…

Artificial Intelligence · Computer Science 2026-05-11 Yi Liu , TingFeng Hui , Wei Zhang , Li Sun , Ningxin Su , Jian Wang , Sen Su

Large Language Model (LLM) simulations, where LLMs act as students with varying approaches to learning tasks, can support teachers' noticing of student thinking. However, simulations using zero- or few-shot prompting often yield inauthentic…

Human-Computer Interaction · Computer Science 2026-04-07 Jie Cao , Ha Nguyen , Selim Yavuz , Boran Yu , Shuguang Wang , Pavneet Kaur Bharaj , Dionne Cross Francis