English
Related papers

Related papers: CausalScore: An Automatic Reference-Free Metric fo…

200 papers

Automatic dialogue coherence evaluation has attracted increasing attention and is crucial for developing promising dialogue systems. However, existing metrics have two major limitations: (a) they are mostly trained in a simplified two-level…

Computation and Language · Computer Science 2021-07-23 Zheng Ye , Liucun Lu , Lishan Huang , Liang Lin , Xiaodan Liang

Question Generation (QG) aims to automate the task of composing questions for a passage with a set of chosen answers found within the passage. In recent years, the introduction of neural generation models has resulted in substantial…

Computation and Language · Computer Science 2022-11-09 Tianbo Ji , Chenyang Lyu , Gareth Jones , Liting Zhou , Yvette Graham

Open-domain human-computer conversation has been attracting increasing attention over the past few years. However, there does not exist a standard automatic evaluation metric for open-domain dialog systems; researchers usually resort to…

Computation and Language · Computer Science 2017-07-18 Chongyang Tao , Lili Mou , Dongyan Zhao , Rui Yan

While witnessing the exceptional success of machine learning (ML) technologies in many applications, users are starting to notice a critical shortcoming of ML: correlation is a poor substitute for causation. The conventional way to discover…

Machine Learning · Computer Science 2024-09-26 Ahmet Kapkiç , Pratanu Mandal , Shu Wan , Paras Sheth , Abhinav Gorantla , Yoonhyuk Choi , Huan Liu , K. Selçuk Candan

Evaluating the quality of generated text automatically remains a significant challenge. Conventional reference-based metrics have been shown to exhibit relatively weak correlation with human evaluations. Recent research advocates the use of…

Computation and Language · Computer Science 2025-11-25 Xiao Wang , Daniil Larionov , Siwei Wu , Yiqi Liu , Steffen Eger , Nafise Sadat Moosavi , Chenghua Lin

Recently, there has been a growing interest in designing text generation systems from a discourse coherence perspective, e.g., modeling the interdependence between sentences. Still, recent BERT-based evaluation metrics are weak in…

Computation and Language · Computer Science 2023-02-07 Wei Zhao , Michael Strube , Steffen Eger

We introduce CLEAR-3K, a dataset of 3,000 assertion-reasoning questions designed to evaluate whether language models can determine if one statement causally explains another. Each question present an assertion-reason pair and challenge…

Computation and Language · Computer Science 2025-06-23 Naiming Liu , Richard Baraniuk , Shashank Sonkar

Encoder-decoder based neural architectures serve as the basis of state-of-the-art approaches in end-to-end open domain dialog systems. Since most of such systems are trained with a maximum likelihood~(MLE) objective they suffer from issues…

The long-standing one-to-many issue of the open-domain dialogues poses significant challenges for automatic evaluation methods, i.e., there may be multiple suitable responses which differ in semantics for a given conversational context. To…

Computation and Language · Computer Science 2023-06-13 Kun Zhao , Bohao Yang , Chenghua Lin , Wenge Rong , Aline Villavicencio , Xiaohui Cui

As large language models (LLMs) witness increasing deployment in complex, high-stakes decision-making scenarios, it becomes imperative to ground their reasoning in causality rather than spurious correlations. However, strong performance on…

Artificial Intelligence · Computer Science 2026-02-24 Yuzhe Wang , Yaochen Zhu , Jundong Li

This paper aims to quantitatively evaluate the performance of ChatGPT, an interactive large language model, on inter-sentential relations such as temporal relations, causal relations, and discourse relations. Given ChatGPT's promising…

Computation and Language · Computer Science 2024-01-29 Chunkit Chan , Jiayang Cheng , Weiqi Wang , Yuxin Jiang , Tianqing Fang , Xin Liu , Yangqiu Song

How reliably an automatic summarization evaluation metric replicates human judgments of summary quality is quantified by system-level correlations. We identify two ways in which the definition of the system-level correlation is inconsistent…

Computation and Language · Computer Science 2022-04-22 Daniel Deutsch , Rotem Dror , Dan Roth

Multimodal Dialogue Summarization (MDS) is a critical task with wide-ranging applications. To support the development of effective MDS models, robust automatic evaluation methods are essential for reducing both cost and human effort.…

Computation and Language · Computer Science 2025-10-03 Yinhong Liu , Jianfeng He , Hang Su , Ruixue Lian , Yi Nian , Jake Vincent , Srikanth Vishnubhotla , Robinson Piramuthu , Saab Mansour

Automatically evaluating the quality of dialogue responses for unstructured domains is a challenging problem. Unfortunately, existing automatic evaluation metrics are biased and correlate very poorly with human judgements of response…

Computation and Language · Computer Science 2018-01-18 Ryan Lowe , Michael Noseworthy , Iulian V. Serban , Nicolas Angelard-Gontier , Yoshua Bengio , Joelle Pineau

High dialogue engagement is a crucial indicator of an effective conversation. A reliable measure of engagement could help benchmark large language models, enhance the effectiveness of human-computer interactions, or improve personal…

Computation and Language · Computer Science 2026-03-17 Yongkang Guo , Zhihuan Huang , Yuqing Kong

Evaluating open-domain dialogue systems is challenging for reasons such as the one-to-many problem, i.e., many appropriate responses other than just the golden response. As of now, automatic evaluation methods need better consistency with…

Computation and Language · Computer Science 2023-09-19 Zhengliang Shi , Weiwei Sun , Shuo Zhang , Zhen Zhang , Pengjie Ren , Zhaochun Ren

Conversational interviews are commonly used to complement structured surveys by eliciting rich and contextualized responses, which are typically analyzed qualitatively. However, their potential contribution to quantitative measurement…

Human-Computer Interaction · Computer Science 2026-03-13 Peinuan Qin , Jingzhu Chen , Yitian Yang , Han Meng , Zicheng Zhu , Yi-Chieh Lee

Automatically evaluating the quality of dialogue responses for unstructured domains is a challenging problem. ADEM(Lowe et al. 2017) formulated the automatic evaluation of dialogue systems as a learning problem and showed that such a model…

Computation and Language · Computer Science 2019-02-26 Ananya B. Sai , Mithun Das Gupta , Mitesh M. Khapra , Mukundhan Srinivasan

We present metrics for evaluating dialog systems through a psychologically-grounded "human" lens in which conversational agents express a diversity of both states (e.g., emotion) and traits (e.g., personality), just as people do. We present…

Computation and Language · Computer Science 2023-09-19 Salvatore Giorgi , Shreya Havaldar , Farhan Ahmed , Zuhaib Akhtar , Shalaka Vaidya , Gary Pan , Lyle H. Ungar , H. Andrew Schwartz , Joao Sedoc

The majority of automatic metrics for evaluating NLG systems are reference-based. However, the challenge of collecting human annotation results in a lack of reliable references in numerous application scenarios. Despite recent advancements…

Computation and Language · Computer Science 2024-03-22 Shuqian Sheng , Yi Xu , Luoyi Fu , Jiaxin Ding , Lei Zhou , Xinbing Wang , Chenghu Zhou
‹ Prev 1 4 5 6 7 8 10 Next ›