English
Related papers

Related papers: Simulating User Satisfaction for the Evaluation of…

200 papers

How do we communicate with others to achieve our goals? We use our prior experience or advice from others, or construct a candidate utterance by predicting how it will be received. However, our experiences are limited and biased, and…

Artificial Intelligence · Computer Science 2023-11-06 Ryan Liu , Howard Yen , Raja Marjieh , Thomas L. Griffiths , Ranjay Krishna

Clarifying the underlying user information need by asking clarifying questions is an important feature of modern conversational search system. However, evaluation of such systems through answering prompted clarifying questions requires…

Computation and Language · Computer Science 2022-04-21 Ivan Sekulić , Mohammad Aliannejadi , Fabio Crestani

User satisfaction with AI assistants is highly personalized: the same response may satisfy one user but disappoint another depending on what each user expects and what they have asked for before. Existing automatic evaluation methods mostly…

Computation and Language · Computer Science 2026-05-29 Zhefan Wang , Zhiqiang Guo , Weizhi Ma , Min Zhang , Quanjia Yan , Hengliang Luo

Advances in large language models (LLMs) enable many new innovations in education. However, evaluating the effectiveness of new technology requires real students, which is time-consuming and hard to scale up. Therefore, many recent works on…

Computation and Language · Computer Science 2026-05-06 Alexander Scarlatos , Jaewook Lee , Simon Woodhead , Andrew Lan

With the increasing popularity of conversational search, how to evaluate the performance of conversational search systems has become an important question in the IR community. Existing works on conversational search evaluation can mainly be…

Information Retrieval · Computer Science 2024-01-25 Zhumin Chu , Qingyao Ai , Zhihong Wang , Yiqun Liu , Yingye Huang , Rui Zhang , Min Zhang , Shaoping Ma

The overall objective of 'social' dialogue systems is to support engaging, entertaining, and lengthy conversations on a wide variety of topics, including social chit-chat. Apart from raw dialogue data, user-provided ratings are the most…

Computation and Language · Computer Science 2018-11-05 Igor Shalyminov , Ondřej Dušek , Oliver Lemon

As dialogue systems and chatbots increasingly integrate into everyday interactions, the need for efficient and accurate evaluation methods becomes paramount. This study explores the comparative performance of human and AI assessments across…

Computation and Language · Computer Science 2024-09-11 Ike Ebubechukwu , Johane Takeuchi , Antonello Ceravola , Frank Joublin

User simulators are crucial for replicating human interactions with dialogue systems, supporting both collaborative training and automatic evaluation, especially for large language models (LLMs). However, current role-playing methods face…

Computation and Language · Computer Science 2025-07-01 Kuang Wang , Xianfei Li , Shenghao Yang , Li Zhou , Feng Jiang , Haizhou Li

Open-domain human-computer conversation has been attracting increasing attention over the past few years. However, there does not exist a standard automatic evaluation metric for open-domain dialog systems; researchers usually resort to…

Computation and Language · Computer Science 2017-07-18 Chongyang Tao , Lili Mou , Dongyan Zhao , Rui Yan

In a typical customer service chat scenario, customers contact a support center to ask for help or raise complaints, and human agents try to solve the issues. In most cases, at the end of the conversation, agents are asked to write a short…

Computation and Language · Computer Science 2021-11-24 Guy Feigenblat , Chulaka Gunasekara , Benjamin Sznajder , Sachindra Joshi , David Konopnicki , Ranit Aharonov

Evaluating open-domain dialogue systems is difficult due to the diversity of possible correct answers. Automatic metrics such as BLEU correlate weakly with human annotations, resulting in a significant bias across different models and…

Computation and Language · Computer Science 2020-04-02 Nouha Dziri , Ehsan Kamalloo , Kory W. Mathewson , Osmar Zaiane

A significant barrier to progress in data-driven approaches to building dialog systems is the lack of high quality, goal-oriented conversational data. To help satisfy this elementary requirement, we introduce the initial release of the…

How to build and use dialogue data efficiently, and how to deploy models in different domains at scale can be two critical issues in building a task-oriented dialogue system. In this paper, we propose a novel manual-guided dialogue scheme…

Computation and Language · Computer Science 2022-08-17 Ryuichi Takanobu , Hao Zhou , Yankai Lin , Peng Li , Jie Zhou , Minlie Huang

In recommendation dialogs, humans commonly disclose their preference and make recommendations in a friendly manner. However, this is a challenge when developing a sociable recommendation dialog system, due to the lack of dialog dataset…

Computation and Language · Computer Science 2020-10-09 Shirley Anugrah Hayati , Dongyeop Kang , Qingxiaoyang Zhu , Weiyan Shi , Zhou Yu

In current text-based task-oriented dialogue (TOD) systems, user emotion detection (ED) is often overlooked or is typically treated as a separate and independent task, requiring additional training. In contrast, our work demonstrates that…

Computation and Language · Computer Science 2024-07-01 Armand Stricker , Patrick Paroubek

There is a resurgent interest in developing intelligent open-domain dialog systems due to the availability of large amounts of conversational data and the recent progress on neural approaches to conversational AI. Unlike traditional…

Computation and Language · Computer Science 2020-03-02 Minlie Huang , Xiaoyan Zhu , Jianfeng Gao

Recommender systems play a central role in numerous real-life applications, yet evaluating their performance remains a significant challenge due to the gap between offline metrics and online behaviors. Given the scarcity and limits (e.g.,…

Information Retrieval · Computer Science 2025-04-18 Nicolas Bougie , Narimasa Watanabe

User Simulators are one of the major tools that enable offline training of task-oriented dialogue systems. For this task the Agenda-Based User Simulator (ABUS) is often used. The ABUS is based on hand-crafted rules and its output is in…

Computation and Language · Computer Science 2018-05-21 Florian Kreyssig , Inigo Casanueva , Pawel Budzianowski , Milica Gasic

Dialogue assessment plays a critical role in the development of open-domain dialogue systems. Existing work are uncapable of providing an end-to-end and human-epistemic assessment dataset, while they only provide sub-metrics like coherence…

Computation and Language · Computer Science 2023-10-26 Yukun Zhao , Lingyong Yan , Weiwei Sun , Chong Meng , Shuaiqiang Wang , Zhicong Cheng , Zhaochun Ren , Dawei Yin

Accurate and interpretable user satisfaction estimation (USE) is critical for understanding, evaluating, and continuously improving conversational systems. Users express their satisfaction or dissatisfaction with diverse conversational…