English
Related papers

Related papers: Improving Statistical Significance in Human Evalua…

200 papers

Score matching estimators have gained widespread attention in recent years partly because they are free from calculating the integral of normalizing constant, thereby addressing the computational challenges in maximum likelihood estimation…

Machine Learning · Statistics 2024-10-08 Haoqun Cao , Zizhuo Meng , Tianjun Ke , Feng Zhou

LLM-as-a-Judge has been widely adopted as an evaluation method and served as supervised rewards in model training. However, existing benchmarks for LLM-as-a-Judge are mainly relying on human-annotated ground truth, which introduces human…

Computation and Language · Computer Science 2025-12-19 Yuanning Feng , Sinan Wang , Zhengxiang Cheng , Yao Wan , Dongping Chen

SAPA is a domain-independent heuristic forward chaining planner that can handle durative actions, metric resource constraints, and deadline goals. It is designed to be capable of handling the multi-objective nature of metric temporal…

Artificial Intelligence · Computer Science 2011-06-28 M. Do , S. Kambhampati

An overview of current debates and contemporary research devoted to the modeling of decision making processes and their facilitation directs attention to the Analytic Hierarchy Process (AHP). At the core of the AHP are various…

Artificial Intelligence · Computer Science 2020-06-05 Paul Thaddeus Kazibudzki

The explosion of open-sourced models and Question-Answering (QA) datasets emphasizes the importance of automated QA evaluation. We studied the statistics of the existing evaluation metrics for a better understanding of their limitations. By…

Computation and Language · Computer Science 2024-10-15 Yun Joon Soh , Jishen Zhao

Comparative evaluation lies at the heart of science, and determining the accuracy of a computational method is crucial for evaluating its potential as well as for guiding future efforts. However, metrics that are typically used have…

Data Analysis, Statistics and Probability · Physics 2019-07-10 Kiwon Um , Xiangyu Hu , Bing Wang , Nils Thuerey

Developing classification algorithms that are fair with respect to sensitive attributes of the data has become an important problem due to the growing deployment of classification algorithms in various social contexts. Several recent works…

Machine Learning · Computer Science 2020-04-16 L. Elisa Celis , Lingxiao Huang , Vijay Keswani , Nisheeth K. Vishnoi

This paper studies a difference operator for stochastic systems whose specifications are represented by Abstract Probabilistic Automata (APAs). In the case refinement fails between two specifications, the target of this operator is to…

Logic in Computer Science · Computer Science 2015-07-01 Benoît Delahaye , Uli Fahrenberg , Kim G. Larsen , Axel Legay

A primary challenge facing modern scientific research is the limited availability of gold-standard data which can be costly, labor-intensive, or invasive to obtain. With the rapid development of machine learning (ML), scientists can now…

Methodology · Statistics 2024-09-17 Jiacheng Miao , Xinran Miao , Yixuan Wu , Jiwei Zhao , Qiongshi Lu

Despite an increasing reliance on fully-automated algorithmic decision-making in our day-to-day lives, human beings still make highly consequential decisions. As frequently seen in business, healthcare, and public policy, recommendations…

Computers and Society · Computer Science 2021-12-14 Kosuke Imai , Zhichao Jiang , James Greiner , Ryan Halen , Sooahn Shin

As Language Model (LM) capabilities advance, evaluating and supervising them at scale is getting harder for humans. There is hope that other language models can automate both these tasks, which we refer to as ''AI Oversight''. We study how…

Conformal prediction is a distribution-free framework for uncertainty quantification that replaces point predictions with sets, offering marginal coverage guarantees (i.e., ensuring that the prediction sets contain the true label with a…

Machine Learning · Computer Science 2025-02-25 Margarida M. Campos , João Calém , Sophia Sklaviadis , Mário A. T. Figueiredo , André F. T. Martins

Reliable evaluation protocols are of utmost importance for reproducible NLP research. In this work, we show that sometimes neither metric nor conventional human evaluation is sufficient to draw conclusions about system performance. Using…

Computation and Language · Computer Science 2021-01-25 Yevgeniy Puzikov

As artificial intelligence plays an increasingly substantial role in decisions affecting humans and society, the accountability of automated decision systems has been receiving increasing attention from researchers and practitioners.…

Machine Learning · Computer Science 2023-07-04 Furkan Gursoy , Ioannis A. Kakadiaris

The pairwise winning indices, computed in the Stochastic Multicriteria Acceptability Analysis, give the probability with which an alternative is preferred to another taking into account all the instances of the assumed preference model…

Optimization and Control · Mathematics 2022-03-29 Sally Giuseppe Arcidiacono , Salvatore Corrente , Salvatore Greco

Automatic evaluation metrics are central to the development of machine translation systems, yet their robustness under domain shift remains unclear. Most metrics are developed on the Workshop on Machine Translation (WMT) benchmarks, raising…

Computation and Language · Computer Science 2026-04-22 Finn Schmidt , Jan Philip Wahle , Terry Ruas , Bela Gipp

High-quality Machine Translation (MT) evaluation relies heavily on human judgments. Comprehensive error classification methods, such as Multidimensional Quality Metrics (MQM), are expensive as they are time-consuming and can only be done by…

Central to human-aligned AI is understanding the benefits of human-elicited labels over synthetic alternatives. While human soft-labels improve calibration by capturing uncertainty, prior studies conflate these benefits with the implicit…

Machine Learning · Computer Science 2026-05-19 Maja Pavlovic , Silviu Paun , Massimo Poesio

Diagnostic accuracy studies assess sensitivity and specificity of a new index test in relation to an established comparator or the reference standard. The development and selection of the index test is usually assumed to be conducted prior…

Methodology · Statistics 2022-08-30 Max Westphal , Antonia Zapf

Automated Scoring (AS), the natural language processing task of scoring essays and speeches in an educational testing setting, is growing in popularity and being deployed across contexts from government examinations to companies providing…

Computation and Language · Computer Science 2021-11-18 Yaman Kumar Singla , Sriram Krishna , Rajiv Ratn Shah , Changyou Chen
‹ Prev 1 3 4 5 6 7 10 Next ›