English
Related papers

Related papers: Is human scoring the best criteria for summary eva…

200 papers

Automatic n-gram based metrics such as ROUGE are widely used for evaluating generative tasks such as summarization. While these metrics are considered indicative (even if imperfect) of human evaluation for English, their suitability for…

Computation and Language · Computer Science 2025-07-14 Itai Mondshine , Tzuf Paz-Argaman , Reut Tsarfaty

Video summarization is a technique to create a short skim of the original video while preserving the main stories/content. There exists a substantial interest in automatizing this process due to the rapid growth of the available material.…

Computer Vision and Pattern Recognition · Computer Science 2019-04-12 Mayu Otani , Yuta Nakashima , Esa Rahtu , Janne Heikkilä

Comparing the differences in outcomes (that is, in "dependent variables") between two subpopulations is often most informative when comparing outcomes only for individuals from the subpopulations who are similar according to "independent…

Methodology · Statistics 2021-12-20 Mark Tygert

In this work, we addressed the issue of combining linear classifiers using their score functions. The value of the scoring function depends on the distance from the decision boundary. Two score functions have been tested and four different…

Machine Learning · Computer Science 2019-05-24 Pawel Trajdos , Robert Burduk

A desirable goal of scientific management is to introduce, if it exists, a simple and reliable way to measure the scientific excellence of publicly-funded research institutions and universities to serve as a basis for their ranking and…

Applications · Statistics 2014-09-23 O. Mryglod , R. Kenna , Yu. Holovatch , B. Berche

The majority of automatic metrics for evaluating NLG systems are reference-based. However, the challenge of collecting human annotation results in a lack of reliable references in numerous application scenarios. Despite recent advancements…

Computation and Language · Computer Science 2024-03-22 Shuqian Sheng , Yi Xu , Luoyi Fu , Jiaxin Ding , Lei Zhou , Xinbing Wang , Chenghu Zhou

The higher criticism of a family of tests starts with the individual uncorrected p-values of each test. It then requires a procedure for deciding whether the collection of p-values indicates the presence of a real effect and if possible…

This paper describes the DSBA submissions to the Prompting Large Language Models as Explainable Metrics shared task, where systems were submitted to two tracks: small and large summarization tracks. With advanced Large Language Models…

Computation and Language · Computer Science 2023-11-08 Joonghoon Kim , Saeran Park , Kiyoon Jeong , Sangmin Lee , Seung Hun Han , Jiyoon Lee , Pilsung Kang

Despite recent advances, evaluating how well large language models (LLMs) follow user instructions remains an open problem. While evaluation methods of language models have seen a rise in prompt-based approaches, limited work on the…

Computation and Language · Computer Science 2023-10-23 Ondrej Skopek , Rahul Aralikatte , Sian Gooding , Victor Carbune

The task of image captioning has recently been gaining popularity, and with it the complex task of evaluating the quality of image captioning models. In this work, we present the first survey and taxonomy of over 70 different image…

Computation and Language · Computer Science 2025-09-16 Uri Berger , Gabriel Stanovsky , Omri Abend , Lea Frermann

Abstractive dialogue summarization has received increasing attention recently. Despite the fact that most of the current dialogue summarization systems are trained to maximize the likelihood of human-written summaries and have achieved…

Computation and Language · Computer Science 2022-12-21 Jiaao Chen , Mohan Dodda , Diyi Yang

Collaborative filtering systems heavily depend on user feedback expressed in product ratings to select and rank items to recommend. In this study we explore how users value different collaborative explanation styles following the user-based…

Information Retrieval · Computer Science 2018-09-07 Ludovik Coba , Markus Zanker , Laurens Rook , Panagiotis Symeonidis

Like it or not, attempts to evaluate and monitor the quality of academic research have become increasingly prevalent worldwide. Performance reviews range from at the level of individuals, through research groups and departments, to entire…

Physics and Society · Physics 2017-03-31 R. Kenna , O. Mryglod , B. Berche

Persistent homology allows us to create topological summaries of complex data. In order to analyse these statistically, we need to choose a topological summary and a relevant metric space in which this topological summary exists. While…

Algebraic Topology · Mathematics 2019-06-24 Katharine Turner , Gard Spreemann

Although text style transfer has witnessed rapid development in recent years, there is as yet no established standard for evaluation, which is performed using several automatic metrics, lacking the possibility of always resorting to human…

Computation and Language · Computer Science 2022-04-18 Huiyuan Lai , Jiali Mao , Antonio Toral , Malvina Nissim

Ranking by pairwise comparisons has shown improved reliability over ordinal classification. However, as the annotations of pairwise comparisons scale quadratically, this becomes less practical when the dataset is large. We propose a method…

Quantitative Methods · Quantitative Biology 2022-02-11 Ikbeom Jang , Garrison Danley , Ken Chang , Jayashree Kalpathy-Cramer

Algorithmic decision systems are increasingly used in areas such as hiring, school admission, or loan approval. Typically, these systems rely on labeled data for training a classification model. However, in many scenarios, ground-truth…

Machine Learning · Computer Science 2021-07-19 Jakob Schoeffer , Niklas Kuehl , Isabel Valera

There is significant interest in developing evaluation metrics which accurately estimate the quality of generated text without the aid of a human-written reference text, which can be time consuming and expensive to collect or entirely…

Computation and Language · Computer Science 2022-10-25 Daniel Deutsch , Rotem Dror , Dan Roth

Binary observations are often repeated to improve data quality, creating technical replicates. Several scoring methods are commonly used to infer the actual individual state and obtain a probability for each state. The common practice of…

Methodology · Statistics 2025-01-24 Manuela Royer-Carenzi , Hadrien Lorenzo , Pierre Pudlo

Estimating the proportion of signals hidden in a large amount of noise variables is of interest in many scientific inquires. In this paper, we consider realistic but theoretically challenging settings with arbitrary covariance dependence…

Methodology · Statistics 2021-04-12 X. Jessie Jeng