English
Related papers

Related papers: Cross-benchmarking for performance evaluation: loo…

200 papers

With the society's growing adoption of machine learning (ML) and deep learning (DL) for various intelligent solutions, it becomes increasingly imperative to standardize a common set of measures for ML/DL models with large scale open…

Machine Learning · Computer Science 2025-04-24 Yen-Hsiang Chang , Jianhao Pu , Wen-mei Hwu , Jinjun Xiong

The recruitment of new personnel is one of the most essential business processes which affect the quality of human capital within any company. It is highly essential for the companies to ensure the recruitment of right talent to maintain a…

Machine Learning · Computer Science 2015-04-09 Gaurav Singh Thakur , Anubhav Gupta , Sangita Gupta

Big data systems address the challenges of capturing, storing, managing, analyzing, and visualizing big data. Within this context, developing benchmarks to evaluate and compare big data systems has become an active topic for both research…

Performance · Computer Science 2014-02-24 Rui Han , Xiaoyi Lu

The increasing versatility of language models (LMs) has given rise to a new class of benchmarks that comprehensively assess a broad range of capabilities. Such benchmarks are associated with massive computational costs, extending to…

Computation and Language · Computer Science 2024-04-02 Yotam Perlitz , Elron Bandel , Ariel Gera , Ofir Arviv , Liat Ein-Dor , Eyal Shnarch , Noam Slonim , Michal Shmueli-Scheuer , Leshem Choshen

Enterprise AI Assistants are increasingly deployed in domains where accuracy is paramount, making each erroneous output a potentially significant incident. This paper presents a comprehensive framework for monitoring, benchmarking, and…

Artificial Intelligence · Computer Science 2025-04-22 Akash V. Maharaj , David Arbour , Daniel Lee , Uttaran Bhattacharya , Anup Rao , Austin Zane , Avi Feller , Kun Qian , Yunyao Li

This paper describes a generalizable model evaluation method that can be adapted to evaluate AI/ML models across multiple criteria including core scientific principles and more practical outcomes. Emerging from prediction competitions in…

Machine Learning · Computer Science 2024-03-19 Jason L. Harman , Jaelle Scheuerman

Numerous studies have compared machine learning (ML) and discrete choice models (DCMs) in predicting travel demand. However, these studies often lack generalizability as they compare models deterministically without considering contextual…

Machine Learning · Computer Science 2025-03-07 Shenhao Wang , Baichuan Mo , Yunhan Zheng , Stephane Hess , Jinhua Zhao

Benchmarks are essential for unified evaluation and reproducibility. The rapid rise of Artificial Intelligence for Software Engineering (AI4SE) has produced numerous benchmarks for tasks such as code generation and bug repair. However, this…

Software Engineering · Computer Science 2025-12-15 Roham Koohestani , Philippe de Bekker , Begüm Koç , Maliheh Izadi

We give an explicit formulaic algorithm and source code for building long-only benchmark portfolios and then using these benchmarks in long-only market outperformance strategies. The benchmarks (or the corresponding betas) do not involve…

Portfolio Management · Quantitative Finance 2018-08-02 Zura Kakushadze , Willie Yu

Reliability and generalization in deep learning are predominantly studied in the context of image classification. Yet, real-world applications in safety-critical domains involve a broader set of semantic tasks, such as semantic segmentation…

Computer Vision and Pattern Recognition · Computer Science 2025-05-26 Shashank Agnihotri , David Schader , Jonas Jakubassa , Nico Sharei , Simon Kral , Mehmet Ege Kaçar , Ruben Weber , Margret Keuper

Benchmark data sets are a cornerstone of machine learning development and applications, ensuring new methods are robust, reliable and competitive. The relative rarity of benchmark sets in computational science, due to the uniqueness of the…

Machine Learning · Computer Science 2025-07-01 Amanda S Barnard

Benchmarks are necessary for healthcare evaluation, but are not sufficient for predicting deployment performance. Our position is that the evaluation--deployment gap arises not because of poorly designed benchmarks, but from implicit…

Computers and Society · Computer Science 2026-05-22 Naveen Raman , Santiago Cortes-Gomez , Mateo Dulce Rubio , Fei Fang , Bryan Wilder

Planning is central to agents and agentic AI. The ability to plan, e.g., creating travel itineraries within a budget, holds immense potential in both scientific and commercial contexts. Moreover, optimal plans tend to require fewer…

Artificial Intelligence · Computer Science 2025-04-22 Haoming Li , Zhaoliang Chen , Jonathan Zhang , Fei Liu

A randomized trial and an analysis of observational data designed to emulate the trial sample observations separately, but have the same eligibility criteria, collect information on some shared baseline covariates, and compare the effects…

Methodology · Statistics 2022-03-29 Issa J. Dahabreh , Jon A. Steingrimsson , James M. Robins , Miguel A. Hernán

Modern benchmarks such as HELM MMLU account for multiple metrics like accuracy, robustness and efficiency. When trying to turn these metrics into a single ranking, natural aggregation procedures can become incoherent or unstable to changes…

Machine Learning · Computer Science 2026-02-10 Polina Gordienko , Christoph Jansen , Julian Rodemann , Georg Schollmeyer

Reliable and robust evaluation methods are a necessary first step towards developing machine learning models that are themselves robust and reliable. Unfortunately, current evaluation protocols typically used to assess classifiers fail to…

Machine Learning · Computer Science 2025-05-26 Michael W. Spratling

The current work proposes an application of DEA methodology for measurement of technical and allocative efficiency of university research activity. The analysis is based on bibliometric data from the Italian university system for the five…

Digital Libraries · Computer Science 2018-11-06 Giovanni Abramo , Tindaro Cicero , Ciriaco Andrea D'Angelo

Benchmarking plays an important role in the development of novel search algorithms as well as for the assessment and comparison of contemporary algorithmic ideas. This paper presents common principles that need to be taken into account when…

Neural and Evolutionary Computing · Computer Science 2018-10-08 Michael Hellwig , Hans-Georg Beyer

We present Task Bench, a parameterized benchmark designed to explore the performance of parallel and distributed programming systems under a variety of application scenarios. Task Bench lowers the barrier to benchmarking multiple…

In a continuous deployment setting, Function-as-a-Service (FaaS) applications frequently receive updated releases, each of which can cause a performance regression. While continuous benchmarking, i.e., comparing benchmark results of the…

Distributed, Parallel, and Cluster Computing · Computer Science 2024-08-20 Tim C. Rese , Nils Japke , Sebastian Koch , Tobias Pfandzelter , David Bermbach
‹ Prev 1 8 9 10 Next ›