English
Related papers

Related papers: Rho-Perfect: Correlation Ceiling For Subjective Ev…

200 papers

The rapid proliferation of large audio models (LAMs) demands efficient approaches for model comparison, yet comprehensive benchmarks are costly. To fill this gap, we investigate whether minimal subsets can reliably evaluate LAMs while…

Computation and Language · Computer Science 2026-05-04 Woody Haosheng Gan , William Held , Diyi Yang

How reliably an automatic summarization evaluation metric replicates human judgments of summary quality is quantified by system-level correlations. We identify two ways in which the definition of the system-level correlation is inconsistent…

Computation and Language · Computer Science 2022-04-22 Daniel Deutsch , Rotem Dror , Dan Roth

Subjective tests are the gold standard for evaluating speech quality and intelligibility; however, they are time-consuming and expensive. Thus, objective measures that align with human perceptions are crucial. This study evaluates the…

Audio and Speech Processing · Electrical Eng. & Systems 2023-10-11 Hsin-Tien Chiang , Kuo-Hsuan Hung , Szu-Wei Fu , Heng-Cheng Kuo , Ming-Hsueh Tsai , Yu Tsao

The purpose of this paper is to pursue our study of rho-estimators built from i.i.d. observations that we defined in Baraud et al. (2014). For a \rho-estimator based on some model S (which means that the estimator belongs to S) and a true…

Statistics Theory · Mathematics 2017-03-07 Yannick Baraud , Lucien Birgé

Human subjective evaluation is the gold standard to evaluate speech quality optimized for human perception. Perceptual objective metrics serve as a proxy for subjective scores. The conventional and widely used metrics require a reference…

Sound · Computer Science 2021-02-12 Chandan K A Reddy , Vishak Gopal , Ross Cutler

For machine learning perception problems, human-level classification performance is used as an estimate of top algorithm performance. Thus, it is important to understand as precisely as possible the factors that impact human-level…

Machine Learning · Computer Science 2019-08-27 Josiah I. Clark , Caroline A. Clark

The field of information retrieval often works with limited and noisy data in an attempt to classify documents into subjective categories, e.g., relevance, sentiment and controversy. We typically quantify a notion of agreement to understand…

Information Retrieval · Computer Science 2018-06-14 John Foley

A challenge in developing machine learning regression models is that it is difficult to know whether maximal performance has been reached on a particular dataset, or whether further model improvement is possible. In biology this problem is…

Biomolecules · Quantitative Biology 2021-07-28 Gang Li , Jan Zrimec , Boyang Ji , Jun Geng , Johan Larsbrink , Aleksej Zelezniak , Jens Nielsen , Martin KM Engqvist

This paper provides a unified framework for analyzing tensor estimation problems that allow for nonlinear observations, heteroskedastic noise, and covariate information. We study a general class of high-dimensional models where each…

Information Theory · Computer Science 2025-06-10 Riccardo Rossetti , Galen Reeves

The quality of human voice plays an important role across various fields like music, speech therapy, and communication, yet it lacks a universally accepted, objective definition. Instead, voice quality is referred to using subjective…

Sound · Computer Science 2024-10-15 Hira Dhamyal , Rita Singh

In the setting of high-dimensional linear models with Gaussian noise, we investigate the possibility of confidence statements connected to model selection. Although there exist numerous procedures for adaptive point estimation, the…

Statistics Theory · Mathematics 2009-10-07 Angelika Rohde , Lutz Duembgen

Objective estimators of multimedia quality are often judged by comparing estimates with subjective "truth data," most often via Pearson correlation coefficient (PCC) or mean-squared error (MSE). But subjective test results contain noise, so…

Audio and Speech Processing · Electrical Eng. & Systems 2026-03-16 Jaden Pieper , Stephen D. Voran

Human feedback has become the de facto standard for evaluating the performance of Large Language Models, and is increasingly being used as a training objective. However, it is not clear which properties of a generated output this single…

Computation and Language · Computer Science 2024-01-17 Tom Hosking , Phil Blunsom , Max Bartolo

This note addresses the question of optimally estimating a linear functional of an object acquired through linear observations corrupted by random noise, where optimality pertains to a worst-case setting tied to a symmetric, convex, and…

Statistics Theory · Mathematics 2023-08-01 Simon Foucart , Grigoris Paouris

The notion of experiment precision quantifies the variance of user ratings in a subjective experiment. Although there exist measures that assess subjective experiment precision, there are no systematic analyses of these measures available…

Multimedia · Computer Science 2022-08-05 Jakub Nawała , Tobias Hoßfeld , Lucjan Janowski , Michael Seufert

When inferring reward functions from human behavior (be it demonstrations, comparisons, physical corrections, or e-stops), it has proven useful to model the human as making noisy-rational choices, with a "rationality coefficient" capturing…

Machine Learning · Computer Science 2023-03-10 Gaurav R. Ghosal , Matthew Zurek , Daniel S. Brown , Anca D. Dragan

This paper discusses two existing approaches to the correlation analysis between automatic evaluation metrics and human scores in the area of natural language generation. Our experiments show that depending on the usage of a system- or…

Computation and Language · Computer Science 2021-03-16 Anastasia Shimorina

In this work, we present the misspecified Gaussian Cram\'er-Rao lower bound for the parameters of a harmonic signal, or pitch, when signal measurements are collected from an almost, but not quite, harmonic model. For the asymptotic case of…

Signal Processing · Electrical Eng. & Systems 2019-10-29 Filip Elvander , Jie Ding , Andreas Jakobsson

We introduce a toolkit for uncovering spurious correlations between recording characteristics and target class in speech datasets. Spurious correlations may arise due to heterogeneous recording conditions, a common scenario for…

Non-reference speech quality models are important for a growing number of applications. The VoiceMOS 2022 challenge provided a dataset of synthetic voice conversion and text-to-speech samples with subjective labels. This study looks at the…

Sound · Computer Science 2022-09-15 Michael Chinen , Jan Skoglund , Chandan K A Reddy , Alessandro Ragano , Andrew Hines
‹ Prev 1 2 3 10 Next ›