English
Related papers

Related papers: On Meta-Evaluation

200 papers

Research in AI evaluation has grown increasingly complex and multidisciplinary, attracting researchers with diverse backgrounds and objectives. As a result, divergent evaluation paradigms have emerged, often developing in isolation,…

Artificial Intelligence · Computer Science 2025-06-09 John Burden , Marko Tešić , Lorenzo Pacchiardi , José Hernández-Orallo

Benchmarks play a significant role in how technology companies communicate about model capabilities and how researchers and the public understand generative AI systems. However, existing benchmarks have been criticized for their failure to…

Human-Computer Interaction · Computer Science 2026-04-29 Charlotte Li , Nick Hagar , Sachita Nishal , Jeremy Gilbert , Nick Diakopoulos

Perceiving and generating diverse modalities are crucial for AI models to effectively learn from and engage with real-world signals, necessitating reliable evaluations for their development. We identify two major issues in current…

Recent advances have made long-form report-generating systems widely available. This has prompted evaluation frameworks that use LLM-as-judge protocols and claim verification, along with meta-evaluation frameworks that seek to validate…

Evaluations of generative models are now ubiquitous, and their outcomes critically shape public and scientific expectations of AI's capabilities. Yet skepticism about their reliability continues to grow. How can we know that a reported…

Artificial Intelligence · Computer Science 2026-05-19 Nathanael Jo , Ashia Wilson

We present MEGA-Bench, an evaluation suite that scales multimodal evaluation to over 500 real-world tasks, to address the highly heterogeneous daily use cases of end users. Our objective is to optimize for a set of high-quality data samples…

Computer Vision and Pattern Recognition · Computer Science 2025-07-15 Jiacheng Chen , Tianhao Liang , Sherman Siu , Zhengqing Wang , Kai Wang , Yubo Wang , Yuansheng Ni , Wang Zhu , Ziyan Jiang , Bohan Lyu , Dongfu Jiang , Xuan He , Yuan Liu , Hexiang Hu , Xiang Yue , Wenhu Chen

Confronted with the challenge of identifying the most suitable metric to validate the merits of newly proposed models, the decision-making process is anything but straightforward. Given that comparing rankings introduces its own set of…

Information Retrieval · Computer Science 2024-08-30 Chiara Balestra , Andreas Mayr , Emmanuel Müller

EXplainable Artificial Intelligence (XAI) aims to help users to grasp the reasoning behind the predictions of an Artificial Intelligence (AI) system. Many XAI approaches have emerged in recent years. Consequently, a subfield related to the…

The rising popularity of explainable artificial intelligence (XAI) to understand high-performing black boxes raised the question of how to evaluate explanations of machine learning (ML) models. While interpretability and explainability are…

Innovations across science and industry are evaluated using randomized trials (a.k.a. A/B tests). While simple and robust, such static designs are inefficient or infeasible for testing many hypotheses. Adaptive designs can greatly improve…

Machine Learning · Computer Science 2024-08-09 Jimmy Wang , Ethan Che , Daniel R. Jiang , Hongseok Namkoong

I would like to share recommendations on how to do performance benchmarks for the purpose of computer science research evaluation. Research in my field (programming language research) often involves performance considerations, but it is…

Programming Languages · Computer Science 2026-05-05 Gabriel Scherer

Most existing emotion analysis emphasizes which emotion arises (e.g., happy, sad, angry) but neglects the deeper why. We propose Emotion Interpretation (EI), focusing on causal factors-whether explicit (e.g., observable objects,…

Artificial Intelligence · Computer Science 2025-04-18 Yuxiang Lin , Jingdong Sun , Zhi-Qi Cheng , Jue Wang , Haomin Liang , Zebang Cheng , Yifei Dong , Jun-Yan He , Xiaojiang Peng , Xian-Sheng Hua

In the rapidly evolving field of artificial intelligence (AI), traditional benchmarks can fall short in attempting to capture the nuanced capabilities of AI models. We focus on the case of physical world modeling and propose a novel…

Artificial Intelligence · Computer Science 2025-09-08 Sasha Mitts

Recent progress in deep research systems has been impressive, but evaluation still lags behind real user needs. Existing benchmarks predominantly assess final reports using fixed rubrics, failing to evaluate the underlying research process.…

Despite significant progress, evaluation of explainable artificial intelligence remains elusive and challenging. In this paper we propose a fine-grained validation framework that is not overly reliant on any one facet of these…

Human-Computer Interaction · Computer Science 2024-03-20 Kacper Sokol , Julia E. Vogt

Scientific evidence often spans instruments, databases, and disciplines, so no single source records the full phenomenon. This makes it difficult to determine when coordinated AI agents add value over simpler scientific workflows. We…

Artificial Intelligence · Computer Science 2026-05-22 Fiona Y. Wong , Markus J. Buehler

This article introduces the Multidimensional Research Assessment Matrix of scientific output. Its base notion holds that the choice of metrics to be applied in a research assessment process depends upon the unit of assessment, the research…

Digital Libraries · Computer Science 2014-06-24 Henk F. Moed , Gali Halevi

Research on Artificial Intelligence (AI)-based Data Assimilation (DA) is expanding rapidly. However, the absence of an objective, comprehensive, and real-world benchmark hinders the fair comparison of diverse methods. Here, we introduce…

Machine Learning · Computer Science 2026-02-17 Wuxin Wang , Weicheng Ni , Ben Fei , Tao Han , Lilan Huang , Taikang Yuan , Xiaoyong Li , Lei Bai , Boheng Duan , Kaijun Ren

Despite significant progress in designing powerful adversarial evasion attacks for robustness verification, the evaluation of these methods often remains inconsistent and unreliable. Many assessments rely on mismatched models, unverified…

Cryptography and Security · Computer Science 2025-07-08 Antonio Emanuele Cinà , Maura Pintor , Luca Demetrio , Ambra Demontis , Battista Biggio , Fabio Roli

The rapid growth of machine learning has produced an ever-expanding ecosystem of models, making it increasingly challenging to verify the reliability of newly released models on unseen, unlabeled data. Conventional evaluation pipelines…

Machine Learning · Computer Science 2026-05-25 Trinh Pham , Viet Huynh , Hongzhi Yin , Quoc Viet Hung Nguyen , Thanh Tam Nguyen