English
Related papers

Related papers: Metrics also Disagree in the Low Scoring Range: Re…

200 papers

Opinion summarization is the automatic creation of text reflecting subjective information expressed in multiple documents, such as user reviews of a product. The task is practically important and has attracted a lot of attention. However,…

Machine Learning · Computer Science 2020-10-13 Arthur Bražinskas , Mirella Lapata , Ivan Titov

Output length is critical to dialogue summarization systems. The dialogue summary length is determined by multiple factors, including dialogue complexity, summary objective, and personal preferences. In this work, we approach dialogue…

Computation and Language · Computer Science 2022-10-28 Bin Wang , Chen Zhang , Chengwei Wei , Haizhou Li

Understanding the quality of a performance evaluation metric is crucial for ensuring that model outputs align with human preferences. However, it remains unclear how well each metric captures the diverse aspects of these preferences, as…

Computation and Language · Computer Science 2025-03-04 Genta Indra Winata , David Anugraha , Lucky Susanto , Garry Kuwanto , Derry Tanti Wijaya

There has been a growing interest in developing machine learning (ML) models for code summarization tasks, e.g., comment generation and method naming. Despite substantial increase in the effectiveness of ML models, the evaluation…

Software Engineering · Computer Science 2022-04-06 Pengyu Nie , Jiyang Zhang , Junyi Jessy Li , Raymond J. Mooney , Milos Gligoric

Text-image generation has advanced rapidly, but assessing whether outputs truly capture the objects, attributes, and relations described in prompts remains a central challenge. Evaluation in this space relies heavily on automated metrics,…

Computer Vision and Pattern Recognition · Computer Science 2025-11-11 Seyed Amir Kasaei , Ali Aghayari , Arash Marioriyad , Niki Sepasian , MohammadAmin Fazli , Mahdieh Soleymani Baghshah , Mohammad Hossein Rohban

A key trait of stochastic optimizers is that multiple runs of the same optimizer in attempting to solve the same problem can produce different results. As a result, their performance is evaluated over several repeats, or runs, on the…

Machine Learning · Computer Science 2026-05-18 Moslem Noori , Elisabetta Valiante , Thomas Van Vaerenbergh , Masoud Mohseni , Ignacio Rozada

To automatically produce a brief yet expressive summary of a long video, an automatic algorithm should start by resembling the human process of summary generation. Prior work proposed supervised and unsupervised algorithms to train models…

Computer Vision and Pattern Recognition · Computer Science 2019-03-04 Mohamed Elfeki , Ali Borji

With the growth of interpreting technologies, from remote interpreting and Computer-Aided Interpreting to automated speech translation and interpreting avatars, there is now a high demand for ways to quickly and efficiently measure the…

Computation and Language · Computer Science 2026-01-12 Jonathan Downie , Joss Moorkens

Summary assessment involves evaluating how well a generated summary reflects the key ideas and meaning of the source text, requiring a deep understanding of the content. Large Language Models (LLMs) have been used to automate this process,…

Computation and Language · Computer Science 2025-12-23 Zahra Sadeghi , Evangelos Milios , Frank Rudzicz

Text simplification intends to make a text easier to read while preserving its core meaning. Intuitively and as shown in previous works, these two dimensions (simplification and meaning preservation) are often-times inversely correlated. An…

Computation and Language · Computer Science 2024-04-05 Liam Cripwell , Joël Legrand , Claire Gardent

Ideal summarization models should generalize to novel summary-worthy content without remembering reference training summaries by rote. However, a single average performance score on the entire test set is inadequate in determining such…

Computation and Language · Computer Science 2023-11-17 Prafulla Kumar Choubey , Alexander R. Fabbri , Caiming Xiong , Chien-Sheng Wu

Automated short-answer scoring lags other LLM applications. We meta-analyze 890 culminating results across a systematic review of LLM short-answer scoring studies, modeling the traditional effect size of Quadratic Weighted Kappa (QWK) with…

Computation and Language · Computer Science 2026-03-27 Michael Hardy

In recent years, various methods and benchmarks have been proposed to empirically evaluate the alignment of artificial neural networks to human neural and behavioral data. But how aligned are different alignment metrics? To answer this…

Neurons and Cognition · Quantitative Biology 2024-07-11 Jannis Ahlert , Thomas Klein , Felix Wichmann , Robert Geirhos

Lack of factual correctness is an issue that still plagues state-of-the-art summarization systems despite their impressive progress on generating seemingly fluent summaries. In this paper, we show that factual inconsistency can be caused by…

Computation and Language · Computer Science 2024-01-22 Asish Ghoshal , Arash Einolghozati , Ankit Arun , Haoran Li , Lili Yu , Vera Gor , Yashar Mehdad , Scott Wen-tau Yih , Asli Celikyilmaz

Automatic text summarization has experienced substantial progress in recent years. With this progress, the question has arisen whether the types of summaries that are typically generated by automatic summarization models align with users'…

Computation and Language · Computer Science 2022-04-26 Maartje ter Hoeve , Julia Kiseleva , Maarten de Rijke

Human evaluation is the foundation upon which the evaluation of both summarization systems and automatic metrics rests. However, existing human evaluation studies for summarization either exhibit a low inter-annotator agreement or have…

Computation and Language · Computer Science 2023-06-07 Yixin Liu , Alexander R. Fabbri , Pengfei Liu , Yilun Zhao , Linyong Nan , Ruilin Han , Simeng Han , Shafiq Joty , Chien-Sheng Wu , Caiming Xiong , Dragomir Radev

Automatic metrics for evaluating translation quality are typically validated by measuring how well they correlate with human assessments. However, correlation methods tend to capture only the ability of metrics to differentiate between good…

Computation and Language · Computer Science 2024-10-11 Sweta Agrawal , António Farinhas , Ricardo Rei , André F. T. Martins

Explainability is widely regarded as essential for trustworthy artificial intelligence systems. However, the metrics commonly used to evaluate counterfactual explanations are algorithmic evaluation metrics that are rarely validated against…

Artificial Intelligence · Computer Science 2026-03-17 Felix Liedeker , Basil Ell , Philipp Cimiano , Christoph Düsing

Automatic machine translation metrics typically rely on human translations to determine the quality of system translations. Common wisdom in the field dictates that the human references should be of very high quality. However, there are no…

Computation and Language · Computer Science 2024-04-11 Vilém Zouhar , Ondřej Bojar

Text summarization, a key natural language generation (NLG) task, is vital in various domains. However, the high cost of inaccurate summaries in risk-critical applications, particularly those involving human-in-the-loop decision-making,…

Computation and Language · Computer Science 2024-10-10 Jianfeng He , Runing Yang , Linlin Yu , Changbin Li , Ruoxi Jia , Feng Chen , Ming Jin , Chang-Tien Lu