English
Related papers

Related papers: Towards Best Experiment Design for Evaluating Dial…

200 papers

Behavioral interview evaluation using large language models presents unique challenges that require structured assessment, realistic interviewer behavior simulation, and pedagogical value for candidate training. We investigate chain of…

Computation and Language · Computer Science 2026-03-12 Kewen Zhu , Zixi Liu , Yanjing Li

Test-time scaling has significantly improved how AI models solve problems, yet current methods often get stuck in repetitive, incorrect patterns of thought. We introduce HEART, a framework that uses emotional cues to guide the model's…

Computation and Language · Computer Science 2026-02-24 Gabriela Pinto , Palash Goyal , Mihir Parmar , Yiwen Song , Souradip Chakraborty , Zifeng Wang , Jinsung Yoon , Tomas Pfister , Hamid Palangi

This paper introduces an adversarial method to stress-test trained metrics to evaluate conversational dialogue systems. The method leverages Reinforcement Learning to find response strategies that elicit optimal scores from the trained…

Artificial Intelligence · Computer Science 2022-03-01 Jan Deriu , Don Tuggener , Pius von Däniken , Mark Cieliebak

Designing service systems requires selecting among alternative configurations -- choosing the best chatbot variant, the optimal routing policy, or the most effective quality control procedure. In many service systems, the primary evidence…

Machine Learning · Computer Science 2026-03-12 Ruicheng Ao , Hongyu Chen , Siyang Gao , Hanwei Li , David Simchi-Levi

Humans follow criteria when they execute tasks, and these criteria are directly used to assess the quality of task completion. Therefore, having models learn to use criteria to provide feedback can help humans or models to perform tasks…

Computation and Language · Computer Science 2024-06-05 Weizhe Yuan , Pengfei Liu , Matthias Gallé

In this paper, we investigate the effect of context on usability evaluation. The focus is on how children behave and perform when they are tested in different settings. Two most commonly applied usability evaluation methods: the think-aloud…

Human-Computer Interaction · Computer Science 2013-06-19 Mohammadi Akheela Khanum , Munesh C. Trivedi

Task-oriented dialogue systems (TODS) are continuing to rise in popularity as various industries find ways to effectively harness their capabilities, saving both time and money. However, even state-of-the-art TODS are not yet reaching their…

Computation and Language · Computer Science 2022-09-07 Ryan Fellows , Hisham Ihshaish , Steve Battle , Ciaran Haines , Peter Mayhew , J. Ignacio Deza

The rise of large language models (LLMs) has brought a critical need for high-quality human-labeled data, particularly for processes like human feedback and evaluation. A common practice is to label data via consensus annotation over human…

Computation and Language · Computer Science 2025-06-23 Manya Wadhwa , Jifan Chen , Junyi Jessy Li , Greg Durrett

Full-Duplex Speech-to-Speech Large Language Models (LLMs) are foundational to natural human-computer interaction, enabling real-time spoken dialogue systems. However, benchmarking and modeling these models remains a fundamental challenge.…

Computation and Language · Computer Science 2025-09-29 Yuan Ge , Saihan Chen , Jingqi Xiao , Xiaoqian Liu , Tong Xiao , Yan Xiang , Zhengtao Yu , Jingbo Zhu

In this position paper, we argue that human evaluation of generative large language models (LLMs) should be a multidisciplinary undertaking that draws upon insights from disciplines such as user experience research and human behavioral…

Computation and Language · Computer Science 2024-10-08 Aparna Elangovan , Ling Liu , Lei Xu , Sravan Bodapati , Dan Roth

This review gives an extensive overview of evaluation methods for task-oriented dialogue systems, paying special attention to practical applications of dialogue systems, for example for customer service. The review (1) provides an overview…

Computation and Language · Computer Science 2024-04-09 Anouck Braggaar , Christine Liebrecht , Emiel van Miltenburg , Emiel Krahmer

Human evaluation is increasingly critical for assessing large language models, capturing linguistic nuances, and reflecting user preferences more accurately than traditional automated metrics. However, the resource-intensive nature of this…

Computation and Language · Computer Science 2023-10-24 Meriem Boubdir , Edward Kim , Beyza Ermis , Marzieh Fadaee , Sara Hooker

This article presents two studies conducted with an affective dialogue system in which text-based system-user communication was used to model, generate, and present different affective and social interaction scenarios. We specifically…

Human-Computer Interaction · Computer Science 2014-06-09 Marcin Skowron , Stefan Rank , Aleksandra Świderska , Dennis Küster , Arvid Kappas

Despite the recent success of automatic metrics for assessing translation quality, their application in evaluating the quality of machine-translated chats has been limited. Unlike more structured texts like news, chat conversations are…

Computation and Language · Computer Science 2024-03-14 Sweta Agrawal , Amin Farajian , Patrick Fernandes , Ricardo Rei , André F. T. Martins

Building a reliable and automated evaluation metric is a necessary but challenging problem for open-domain dialogue systems. Recent studies proposed evaluation metrics that assess generated responses by considering their relevance to…

Computation and Language · Computer Science 2024-07-19 ChaeHun Park , Minseok Choi , Dohyun Lee , Jaegul Choo

Rating scales are a widely used method for data annotation; however, they present several challenges, such as difficulty in maintaining inter- and intra-annotator consistency. Best-worst scaling (BWS) is an alternative method of annotation…

Computation and Language · Computer Science 2017-12-06 Svetlana Kiritchenko , Saif M. Mohammad

How reliably an automatic summarization evaluation metric replicates human judgments of summary quality is quantified by system-level correlations. We identify two ways in which the definition of the system-level correlation is inconsistent…

Computation and Language · Computer Science 2022-04-22 Daniel Deutsch , Rotem Dror , Dan Roth

Automatic evaluation of open-domain dialogue response generation is very challenging because there are many appropriate responses for a given context. Existing evaluation models merely compare the generated response with the ground truth…

Computation and Language · Computer Science 2020-06-15 JinYeong Bak , Alice Oh

Open-domain human-computer conversation has been attracting increasing attention over the past few years. However, there does not exist a standard automatic evaluation metric for open-domain dialog systems; researchers usually resort to…

Computation and Language · Computer Science 2017-07-18 Chongyang Tao , Lili Mou , Dongyan Zhao , Rui Yan

Recent dialogue coherence models use the coherence features designed for monologue texts, e.g. nominal entities, to represent utterances and then explicitly augment them with dialogue-relevant features, e.g., dialogue act labels. It…

Computation and Language · Computer Science 2020-06-04 Mohsen Mesgar , Sebastian Bücker , Iryna Gurevych