English
Related papers

Related papers: On Meta-Evaluation

200 papers

Evaluation plays a crucial role in the advancement of information retrieval (IR) models. However, current benchmarks, which are based on predefined domains and human-labeled data, face limitations in addressing evaluation needs for emerging…

Information Retrieval · Computer Science 2025-07-25 Jianlyu Chen , Nan Wang , Chaofan Li , Bo Wang , Shitao Xiao , Han Xiao , Hao Liao , Defu Lian , Zheng Liu

In the field of modeling, the word validation refers to simple comparisons between model outputs and experimental data. Usually, this comparison constitutes plotting the model results against data on the same axes to provide a visual…

Applications · Statistics 2021-06-11 Farid Mohammadi

Comparative evaluation lies at the heart of science, and determining the accuracy of a computational method is crucial for evaluating its potential as well as for guiding future efforts. However, metrics that are typically used have…

Data Analysis, Statistics and Probability · Physics 2019-07-10 Kiwon Um , Xiangyu Hu , Bing Wang , Nils Thuerey

The AEC-Bench is a multimodal benchmark for evaluating agentic systems on real-world tasks in the Architecture, Engineering, and Construction (AEC) domain. The benchmark covers tasks requiring drawing understanding, cross-sheet reasoning,…

Artificial Intelligence · Computer Science 2026-04-01 Harsh Mankodiya , Chase Gallik , Theodoros Galanos , Andriy Mulyar

Evidence-based education has become a central concept in science education, with meta-analyses often regarded as the gold standard for informing practice. This emphasis raises critical questions concerning the applicability,…

Physics Education · Physics 2026-02-06 Christoph Kulgemeyer , Anna Weißbach , Kasim Costan , David Geelan , David Treagust

Automated feedback systems have become increasingly integral to programming education, where learners engage in iterative cycles of code construction, testing, and refinement. Despite its wider integration in practices and technical…

Computers and Society · Computer Science 2026-02-03 Yeonji Jung , Yunseo Lee , Jiyeong Bae , DoYong Kim , Heungsoo Choi , Minji Kang , Unggi Lee

In NLG meta-evaluation, evaluation metrics are typically assessed based on their consistency with humans. However, we identify some limitations in traditional NLG meta-evaluation approaches, such as issues in handling human ratings and…

Computation and Language · Computer Science 2025-08-18 Xinyu Hu , Mingqi Gao , Li Lin , Zhenghan Yu , Xiaojun Wan

This paper systematically derives design dimensions for the structured evaluation of explainable artificial intelligence (XAI) approaches. These dimensions enable a descriptive characterization, facilitating comparisons between different…

Human-Computer Interaction · Computer Science 2020-09-15 Fabian Sperrle , Mennatallah El-Assady , Grace Guo , Duen Horng Chau , Alex Endert , Daniel Keim

The objective of this paper is to design performance metrics and respective formulas to quantitatively evaluate the achievement of set objectives and expected outcomes both at the course and program levels. Evaluation is defined as one or…

Physics Education · Physics 2015-09-16 Irfan Ahmed , Arif Bhatti

Evaluation of potential AGI systems and methods is difficult due to the breadth of the engineering goal. We have no methods for perfect evaluation of the end state, and instead measure performance on small tests designed to provide…

Artificial Intelligence · Computer Science 2025-10-03 John Hawkins

Forming a reliable judgement of a machine learning (ML) model's appropriateness for an application ecosystem is critical for its responsible use, and requires considering a broad range of factors including harms, benefits, and…

Machine Learning · Computer Science 2022-05-12 Ben Hutchinson , Negar Rostamzadeh , Christina Greer , Katherine Heller , Vinodkumar Prabhakaran

How can one meaningfully make a measurement, if the meter does not conform to any standard and its scale expands or shrinks depending on what is measured? In the present work it is argued that current evaluation practices for…

Machine Learning · Computer Science 2023-02-24 K. Dyrland , A. S. Lundervold , P. G. L. Porta Mana

Learned representations of scientific documents can serve as valuable input features for downstream tasks without further fine-tuning. However, existing benchmarks for evaluating these representations fail to capture the diversity of…

Computation and Language · Computer Science 2023-11-14 Amanpreet Singh , Mike D'Arcy , Arman Cohan , Doug Downey , Sergey Feldman

Experimentation is an intrinsic part of research in artificial intelligence since it allows for collecting quantitative observations, validating hypotheses, and providing evidence for their reformulation. For that reason, experimentation…

Artificial Intelligence · Computer Science 2024-02-14 Josu Ceberio , Borja Calvo

We address a fundamental challenge in Natural Language Generation (NLG) model evaluation -- the design and evaluation of evaluation metrics. Recognizing the limitations of existing automatic metrics and noises from how current human…

Computation and Language · Computer Science 2023-10-24 Ziang Xiao , Susu Zhang , Vivian Lai , Q. Vera Liao

The meteoric rise of AI, with its rapidly expanding market capitalization, presents both transformative opportunities and critical challenges. Chief among these is the urgent need for a new, unified paradigm for trustworthy evaluation, as…

In empirical software engineering, benchmarks can be used for comparing different methods, techniques and tools. However, the recent ACM SIGSOFT Empirical Standards for Software Engineering Research do not include an explicit checklist for…

Software Engineering · Computer Science 2021-05-04 Wilhelm Hasselbring

Scientific evaluation is a determinant of how scientists, institutions and funders behave, and as such is a key element in the making of science. In this article, we propose an alternative to the current norm of evaluating research with…

Digital Libraries · Computer Science 2017-01-30 Michaël Bon , Michael Taylor , Gary S. McDowell

Recent advances in deep research systems enable large language models to retrieve, synthesize, and reason over large-scale external knowledge. In medicine, developing clinical guidelines critically depends on such deep evidence integration.…

Reliability has long been treated as an engineering practice supported by testing, statistics and standards, yet its status as a scientific discipline remains unsettled. From a philosophical perspective, scientific truth is characterized by…

Physics and Society · Physics 2026-01-13 Xiao-Yang Li , Shi-Shun Chen , Waichon Lio , Rui Kang
‹ Prev 1 3 4 5 6 7 10 Next ›