English
Related papers

Related papers: Exploring the Impact of Human Evaluator Group on C…

200 papers

Automatic evaluation of open-domain dialogs remains an unsolved problem. Moreover, existing methods do not correlate strongly with human annotations. This paper presents a new automated evaluation method using follow-ups: we measure the…

Computation and Language · Computer Science 2022-09-13 Maxime De Bruyn , Ehsan Lotfi , Jeska Buhmann , Walter Daelemans

In the realm of human-AI dialogue, the facilitation of empathetic responses is important. Validation is one of the key communication techniques in psychology, which entails recognizing, understanding, and acknowledging others' emotional…

Computation and Language · Computer Science 2024-02-21 Zi Haur Pang , Yahui Fu , Divesh Lala , Keiko Ochi , Koji Inoue , Tatsuya Kawahara

Despite significant research effort in the development of automatic dialogue evaluation metrics, little thought is given to evaluating dialogues other than in English. At the same time, ensuring metrics are invariant to semantically similar…

Computation and Language · Computer Science 2023-09-11 John Mendonça , Patrícia Pereira , Helena Moniz , João Paulo Carvalho , Alon Lavie , Isabel Trancoso

We describe and validate a metric for estimating multi-class classifier performance based on cross-validation and adapted for improvement of small, unbalanced natural-language datasets used in chatbot design. Our experiences draw upon…

Information Retrieval · Computer Science 2019-06-06 Kit Kuksenok , Andriy Martyniv

With the rapid adoption of LLM-based chatbots, there is a pressing need to evaluate what humans and LLMs can achieve together. However, standard benchmarks, such as MMLU, measure LLM capabilities in isolation (i.e., "AI-alone"). Here, we…

Computation and Language · Computer Science 2025-08-13 Serina Chang , Ashton Anderson , Jake M. Hofman

Recently, the emergence of ChatGPT has attracted wide attention from the computational linguistics community. Many prior studies have shown that ChatGPT achieves remarkable performance on various NLP tasks in terms of automatic evaluation…

Computation and Language · Computer Science 2023-10-25 Jiaan Wang , Yunlong Liang , Fandong Meng , Zengkui Sun , Haoxiang Shi , Zhixu Li , Jinan Xu , Jianfeng Qu , Jie Zhou

User engagement is a critical metric for evaluating the quality of open-domain dialogue systems. Prior work has focused on conversation-level engagement by using heuristically constructed features such as the number of turns and the total…

Computation and Language · Computer Science 2020-01-27 Sarik Ghazarian , Ralph Weischedel , Aram Galstyan , Nanyun Peng

Dialogue Systems are tools designed for various practical purposes concerning human-machine interaction. These systems should be built on ethical foundations because their behavior may heavily influence a user (think especially about…

Artificial Intelligence · Computer Science 2021-09-20 Abeer Dyoub , Stefania Costantini , Ivan Letteri , Francesca A. Lisi

Current state-of-the-art neural dialogue systems are mainly data-driven and are trained on human-generated responses. However, due to the subjectivity and open-ended nature of human conversations, the complexity of training dialogues varies…

Computation and Language · Computer Science 2020-03-17 Hengyi Cai , Hongshen Chen , Cheng Zhang , Yonghao Song , Xiaofang Zhao , Yangxi Li , Dongsheng Duan , Dawei Yin

In cognitive science and linguistic theory, dialogue is not seen as a chain of independent utterances but rather as a joint activity sustained by coherence, consistency, and shared understanding. However, many systems for open-domain and…

Computation and Language · Computer Science 2026-03-24 Tianyi Zhang , David Traum

Reliable automatic evaluation of dialogue systems under an interactive environment has long been overdue. An ideal environment for evaluating dialog systems, also known as the Turing test, needs to involve human interaction, which is…

Computation and Language · Computer Science 2021-09-23 Haoming Jiang , Bo Dai , Mengjiao Yang , Tuo Zhao , Wei Wei

The advent of large language models (LLMs), such as GPT, Gemini, and DeepSeek, has significantly advanced natural language processing, giving rise to sophisticated chatbots capable of diverse language-related tasks. The transition from…

Computation and Language · Computer Science 2025-06-16 Jiachen Zhu , Menghui Zhu , Renting Rui , Rong Shan , Congmin Zheng , Bo Chen , Yunjia Xi , Jianghao Lin , Weiwen Liu , Ruiming Tang , Yong Yu , Weinan Zhang

In recent years, several high-performance conversational systems have been proposed based on the Transformer encoder-decoder model. Although previous studies analyzed the effects of the model parameters and the decoding method on subjective…

Computation and Language · Computer Science 2021-09-14 Hiroaki Sugiyama , Masahiro Mizukami , Tsunehiro Arimoto , Hiromi Narimatsu , Yuya Chiba , Hideharu Nakajima , Toyomi Meguro

Large language models (LLMs) increasingly operate in environments where they encounter social information such as other agents' answers, tool outputs, or human recommendations. In humans, such inputs influence judgments in ways that depend…

Artificial Intelligence · Computer Science 2026-02-17 Anooshka Bajaj , Zoran Tiganj

As AI chatbots increasingly incorporate empathy, understanding user-centered perceptions of chatbot empathy and its impact on conversation quality remains essential yet under-explored. This study examines how chatbot identity and perceived…

Human-Computer Interaction · Computer Science 2025-03-10 Tingting Liu , Salvatore Giorgi , Ankit Aich , Allison Lahnala , Brenda Curtis , Lyle Ungar , João Sedoc

Context: A deeper understanding of human factors in software engineering (SE) is essential for improving team collaboration, decision-making, and productivity. Communication channels like code reviews and chats provide insights into…

Software Engineering · Computer Science 2025-04-21 Amirali Sajadi , Kostadin Damevski , Preetha Chatterjee

Dialogue systems in the form of chatbots and personal assistants are being increasingly integrated into people's lives. Modern dialogue systems may consider adopting anthropomorphic personas, mimicking societal demographic groups to appear…

Computation and Language · Computer Science 2021-12-16 Emily Sheng , Josh Arnold , Zhou Yu , Kai-Wei Chang , Nanyun Peng

The explosion of high-performing conversational language models (LMs) has spurred a shift from classic natural language processing (NLP) benchmarks to expensive, time-consuming and noisy human evaluations - yet the relationship between…

Evaluation is crucial in the development process of task-oriented dialogue systems. As an evaluation method, user simulation allows us to tackle issues such as scalability and cost-efficiency, making it a viable choice for large-scale…

Information Retrieval · Computer Science 2021-05-11 Weiwei Sun , Shuo Zhang , Krisztian Balog , Zhaochun Ren , Pengjie Ren , Zhumin Chen , Maarten de Rijke

Dialogue systems capable of social influence such as persuasion, negotiation, and therapy, are essential for extending the use of technology to numerous realistic scenarios. However, existing research primarily focuses on either…

Computation and Language · Computer Science 2023-01-26 Kushal Chawla , Weiyan Shi , Jingwen Zhang , Gale Lucas , Zhou Yu , Jonathan Gratch