中文
相关论文

相关论文: Synthetic Users, Real Differences: an Evaluation F…

200 篇论文

Long-term, open-domain dialogue capabilities are essential for chatbots aiming to recall past interactions and demonstrate emotional intelligence (EI). Yet, most existing research relies on synthetic, LLM-generated data, leaving open…

计算与语言 · 计算机科学 2025-02-20 Dong-Ho Lee , Adyasha Maharana , Jay Pujara , Xiang Ren , Francesco Barbieri

Task-oriented conversational systems are essential for efficiently addressing diverse user needs, yet their development requires substantial amounts of high-quality conversational data that is challenging and costly to obtain. While large…

信息检索 · 计算机科学 2025-11-06 Zhefan Wang , Ning Geng , Zhiqiang Guo , Weizhi Ma , Min Zhang

User simulation is a promising approach for automatically training and evaluating conversational information access agents, enabling the generation of synthetic dialogues and facilitating reproducible experiments at scale. However, the…

信息检索 · 计算机科学 2024-06-28 Nolwenn Bernard , Krisztian Balog

Recent advancements in Large Language Models (LLMs) have significantly enhanced conversational agents, making them applicable to various fields (e.g., education, entertainment). Despite their progress, the evaluation of the agents often…

计算与语言 · 计算机科学 2025-09-29 Jiho Kim , Woosog Chay , Hyeonji Hwang , Daeun Kyung , Hyunseung Chung , Eunbyeol Cho , Yeonsu Kwon , Yohan Jo , Edward Choi

Synthetic data is increasingly critical for contact centers, where privacy constraints and data scarcity limit the availability of real conversations. However, generating synthetic dialogues that are realistic and useful for downstream…

计算与语言 · 计算机科学 2026-02-17 Rishikesh Devanathan , Varun Nathan , Ayush Kumar

Simulating real personalities with large language models requires grounding generation in authentic personal data. Existing evaluation approaches rely on demographic surveys, personality questionnaires, or short AI-led interviews as…

计算与语言 · 计算机科学 2026-02-25 Yu Li , Pranav Narayanan Venkit , Yada Pruksachatkun , Chien-Sheng Wu

Conversational recommender systems (CRS) enhance user experience through multi-turn interactions, yet evaluating CRS remains challenging. User simulators can provide comprehensive evaluations through interactions with CRS, but building…

人机交互 · 计算机科学 2025-08-01 Luyu Chen , Quanyu Dai , Zeyu Zhang , Xueyang Feng , Mingyu Zhang , Pengcheng Tang , Xu Chen , Yue Zhu , Zhenhua Dong

Conversational AI chatbots are transforming industries by streamlining customer service, automating transactions, and enhancing user engagement. However, evaluating these systems remains a challenge, particularly in financial services,…

计算机与社会 · 计算机科学 2025-02-11 Shailja Gupta , Rajesh Ranjan , Surya Narayan Singh

Synthetic users are cost-effective proxies for real users in the evaluation of conversational recommender systems. Large language models show promise in simulating human-like behavior, raising the question of their ability to represent a…

计算与语言 · 计算机科学 2024-03-27 Se-eun Yoon , Zhankui He , Jessica Maria Echterhoff , Julian McAuley

Task-oriented dialogue systems (TDSs) are assessed mainly in an offline setting or through human evaluation. The evaluation is often limited to single-turn or is very time-intensive. As an alternative, user simulators that mimic user…

计算与语言 · 计算机科学 2023-11-06 Weiwei Sun , Shuyu Guo , Shuo Zhang , Pengjie Ren , Zhumin Chen , Maarten de Rijke , Zhaochun Ren

Evaluation is crucial in the development process of task-oriented dialogue systems. As an evaluation method, user simulation allows us to tackle issues such as scalability and cost-efficiency, making it a viable choice for large-scale…

信息检索 · 计算机科学 2021-05-11 Weiwei Sun , Shuo Zhang , Krisztian Balog , Zhaochun Ren , Pengjie Ren , Zhumin Chen , Maarten de Rijke

The development of chatbots requires collecting a large number of human-chatbot dialogues to reflect the breadth of users' sociodemographic backgrounds and conversational goals. However, the resource requirements to conduct the respective…

计算与语言 · 计算机科学 2024-10-15 Hovhannes Tamoyan , Hendrik Schuff , Iryna Gurevych

Doctor-patient consultations require multi-turn, context-aware communication tailored to diverse patient personas. Training or evaluating doctor LLMs in such settings requires realistic patient interaction systems. However, existing…

人工智能 · 计算机科学 2025-10-30 Daeun Kyung , Hyunseung Chung , Seongsu Bae , Jiho Kim , Jae Ho Sohn , Taerim Kim , Soo Kyung Kim , Edward Choi

User simulators are essential for training reinforcement learning (RL) based dialog models. The performance of the simulator directly impacts the RL policy. However, building a good user simulator that models real user behaviors is…

计算与语言 · 计算机科学 2019-09-05 Weiyan Shi , Kun Qian , Xuewei Wang , Zhou Yu

LLM-based user simulation is the primary mechanism for end-to-end agent evaluation, yet simulated users are poor proxies for real humans: unconstrained LLM defaults produce a Formalism Ceiling (style match rates of 6-8% against real users),…

Large Language Model (LLM) personas with explicit specifications of attributes, background, and behavioural tendencies are increasingly used to simulate human conversations for tasks such as user modeling, social reasoning, and behavioural…

计算与语言 · 计算机科学 2026-03-04 Eliseo Bao , Anxo Perez , Xi Wang , Javier Parapar

Synthetic therapy dialogues generated by large language models (LLMs) are increasingly used in mental health NLP to simulate counseling scenarios, train models, and supplement limited real-world data. However, it remains unclear whether…

计算与语言 · 计算机科学 2025-12-18 Xiaoyi Wang , Jiwei Zhang , Guangtao Zhang , Honglei Guo

As Large Language Models (LLMs) evolve from static dialogue interfaces to autonomous general agents, effective memory is paramount to ensuring long-term consistency. However, existing benchmarks primarily focus on casual conversation or…

计算与语言 · 计算机科学 2026-01-13 Haonan Bian , Zhiyuan Yao , Sen Hu , Zishan Xu , Shaolei Zhang , Yifu Guo , Ziliang Yang , Xueran Han , Huacan Wang , Ronghao Chen

We present BotSIM, a data-efficient end-to-end Bot SIMulation toolkit for commercial text-based task-oriented dialog (TOD) systems. BotSIM consists of three major components: 1) a Generator that can infer semantic-level dialog acts and…

计算与语言 · 计算机科学 2022-12-01 Guangsen Wang , Samson Tan , Shafiq Joty , Gang Wu , Jimmy Au , Steven Hoi

Conversational information access is an emerging research area. Currently, human evaluation is used for end-to-end system evaluation, which is both very time and resource intensive at scale, and thus becomes a bottleneck of progress. As an…

信息检索 · 计算机科学 2020-06-17 Shuo Zhang , Krisztian Balog
‹ 上一页 1 2 3 10 下一页 ›