English
Related papers

Related papers: Improving Statistical Significance in Human Evalua…

200 papers

While advanced classifiers have been increasingly used in real-world safety-critical applications, how to properly evaluate the black-box models given specific human values remains a concern in the community. Such human values include…

Machine Learning · Computer Science 2024-03-14 Yanyun Wang , Dehui Du , Yuanhao Liu

Recent successes suggest that parameter-efficient fine-tuning of foundation models as the state-of-the-art method for transfer learning in vision, replacing the rich literature of alternatives such as meta-learning. In trying to harness the…

Computer Vision and Pattern Recognition · Computer Science 2024-07-02 Shengzhuang Chen , Jihoon Tack , Yunqiao Yang , Yee Whye Teh , Jonathan Richard Schwarz , Ying Wei

Automatic metrics are commonly used as the exclusive tool for declaring the superiority of one machine translation system's quality over another. The community choice of automatic metric guides research directions and industrial…

Computation and Language · Computer Science 2021-09-15 Tom Kocmi , Christian Federmann , Roman Grundkiewicz , Marcin Junczys-Dowmunt , Hitokazu Matsushita , Arul Menezes

Due to the susceptibility of Artificial Intelligence (AI) to data perturbations and adversarial examples, it is crucial to perform a thorough robustness evaluation before any Machine Learning (ML) model is deployed. However, examining a…

Machine Learning · Computer Science 2025-10-01 João Vitorino , Eva Maia , Isabel Praça , Carlos Soares

Recent advances have made long-form report-generating systems widely available. This has prompted evaluation frameworks that use LLM-as-judge protocols and claim verification, along with meta-evaluation frameworks that seek to validate…

Subjective assessment tests are often employed to evaluate image processing systems, notably image and video compression, super-resolution among others and have been used as an indisputable way to provide evidence of the performance of an…

Multimedia · Computer Science 2023-11-13 Shima Mohammadi , Joao Ascenso

Aligning large language models (LLMs) with human preferences becomes a key component to obtaining state-of-the-art performance, but it yields a huge cost to construct a large human-annotated preference dataset. To tackle this problem, we…

Machine Learning · Computer Science 2025-03-05 Dongyoung Kim , Kimin Lee , Jinwoo Shin , Jaehyung Kim

This paper introduces Pairwise Difference Pearson (PDP), a novel segment-level meta-evaluation metric for Machine Translation (MT) that address limitations in previous Pearson's $\rho$-based and and Kendall's $\tau$-based meta-evaluation…

Computation and Language · Computer Science 2025-10-01 Colten DiIanni , Daniel Deutsch

We present MetaMetrics-MT, an innovative metric designed to evaluate machine translation (MT) tasks by aligning closely with human preferences through Bayesian optimization with Gaussian Processes. MetaMetrics-MT enhances existing MT…

Computation and Language · Computer Science 2024-11-04 David Anugraha , Garry Kuwanto , Lucky Susanto , Derry Tanti Wijaya , Genta Indra Winata

How reliably an automatic summarization evaluation metric replicates human judgments of summary quality is quantified by system-level correlations. We identify two ways in which the definition of the system-level correlation is inconsistent…

Computation and Language · Computer Science 2022-04-22 Daniel Deutsch , Rotem Dror , Dan Roth

Annually, at the Conference of Machine Translation (WMT), the Metrics Shared Task organizers conduct the meta-evaluation of Machine Translation (MT) metrics, ranking them according to their correlation with human judgments. Their results…

Computation and Language · Computer Science 2024-08-27 Stefano Perrella , Lorenzo Proietti , Alessandro Scirè , Edoardo Barba , Roberto Navigli

Large language models (LLMs) are increasingly used as judges to replace costly human preference labels in pairwise evaluation. Despite their practicality, LLM judges remain prone to miscalibration and systematic biases. This paper proposes…

Computation and Language · Computer Science 2026-02-20 Sher Badshah , Ali Emami , Hassan Sajjad

With the increasing complexity of modern industrial automatic and robotic systems, an increasing burden is put on the operators, who are requested to supervise and interact with such complex systems, typically under challenging and…

Human-Computer Interaction · Computer Science 2020-03-06 Lorenzo Sabattini , Valeria Villani , Julia N. Czerniak , Frieder Loch , Alexander Mertens , Birgit Vogel-Heuser , Cesare Fantuzzi

Preference alignment is usually achieved by weight-updating training on preference data, which adds substantial alignment-stage compute and provides limited mechanistic visibility. We propose Dynamic SAE Steering for Preference Alignment…

Machine Learning · Computer Science 2026-03-24 James Wedgwood , Aashiq Muhamed , Mona T. Diab , Virginia Smith

Automated metrics for Machine Translation have made significant progress, with the goal of replacing expensive and time-consuming human evaluations. These metrics are typically assessed by their correlation with human judgments, which…

Computation and Language · Computer Science 2024-12-31 Pius von Däniken , Jan Deriu , Mark Cieliebak

Background: High-throughput proteomics techniques, such as mass spectrometry (MS)-based approaches, produce very high-dimensional data-sets. In a clinical setting one is often interested in how mass spectra differ between patients of…

Annually, research teams spend large amounts of money to evaluate the quality of machine translation systems (WMT, inter alia). This is expensive because it requires a lot of expert human labor. In the recently adopted annotation protocol,…

Computation and Language · Computer Science 2025-01-30 Vilém Zouhar , Tom Kocmi , Mrinmaya Sachan

With the growing access to administrative health databases, retrospective studies have become crucial evidence for medical treatments. Yet, non-randomized studies frequently face selection biases, requiring mitigation strategies. Propensity…

Machine Learning · Statistics 2026-05-07 Alexandre Abraham , Andrés Hoyos Idrobo

When Perturbation Analysis (PA) yields unbiased sensitivity estimators for expected-value performance functions in discrete event dynamic systems, it can be used for performance optimization of those functions. However, when PA is known to…

Optimization and Control · Mathematics 2013-08-06 Yorai Wardi , Christos G. Cassandras

Kendall's tau is frequently used to meta-evaluate how well machine translation (MT) evaluation metrics score individual translations. Its focus on pairwise score comparisons is intuitive but raises the question of how ties should be…

Computation and Language · Computer Science 2023-10-18 Daniel Deutsch , George Foster , Markus Freitag