English
Related papers

Related papers: Two-step benchmarking: Setting more realistically …

200 papers

Motivated by the need for adaptive, secure and responsive scheduling in a great range of computing applications, including human-centered and time-critical applications, this paper proposes a scheduling framework that seamlessly adds…

Distributed, Parallel, and Cluster Computing · Computer Science 2020-01-14 Georgios C. Chasparis , Vladimir Janjic , Michael Rossbory

As large language models (LLMs) become increasingly versatile, numerous large scale benchmarks have been developed to thoroughly assess their capabilities. These benchmarks typically consist of diverse datasets and prompts to evaluate…

Machine Learning · Computer Science 2024-10-10 Yang Li , Jie Ma , Miguel Ballesteros , Yassine Benajiba , Graham Horwood

The deployment of large-scale models, such as large language models (LLMs) and sophisticated image generation systems, incurs substantial costs due to their computational demands. To mitigate these costs and address challenges related to…

Machine Learning · Computer Science 2024-10-30 Yuzhe Yang , Yipeng Du , Ahmad Farhan , Claudio Angione , Yue Zhao , Harry Yang , Fielding Johnston , James Buban , Patrick Colangelo

The increasing versatility of language models (LMs) has given rise to a new class of benchmarks that comprehensively assess a broad range of capabilities. Such benchmarks are associated with massive computational costs, extending to…

Computation and Language · Computer Science 2024-04-02 Yotam Perlitz , Elron Bandel , Ariel Gera , Ofir Arviv , Liat Ein-Dor , Eyal Shnarch , Noam Slonim , Michal Shmueli-Scheuer , Leshem Choshen

In policy learning, the goal is typically to optimize a primary performance metric, but other subsidiary metrics often also warrant attention. This paper presents two strategies for evaluating these subsidiary metrics under a policy that is…

Statistics Theory · Mathematics 2025-09-09 Zhaoqi Li , Houssam Nassif , Alex Luedtke

Evaluating models on large benchmarks is very resource-intensive, especially during the period of rapid model evolution. Existing efficient evaluation methods estimate the performance of target models by testing them only on a small and…

Machine Learning · Computer Science 2025-06-03 Peiwen Yuan , Yueqi Zhang , Shaoxiong Feng , Yiwei Li , Xinglin Wang , Jiayi Shi , Chuyi Tan , Boyuan Pan , Yao Hu , Kan Li

Numerous benchmarks for Few-Shot Learning have been proposed in the last decade. However all of these benchmarks focus on performance averaged over many tasks, and the question of how to reliably evaluate and tune models trained for…

Machine Learning · Computer Science 2023-07-07 Luísa Shimabucoro , Timothy Hospedales , Henry Gouk

The performance model of an application can pro- vide understanding about its runtime behavior on particular hardware. Such information can be analyzed by developers for performance tuning. However, model building and analyzing is…

Performance · Computer Science 2017-05-23 Kewen Meng , Boyana Norris

In order to track and comprehend the academic achievement of students, both private and public educational institutions devote a significant amount of resources and labour. One of the difficult issues that institutes deal with on a regular…

Computers and Society · Computer Science 2022-11-14 Bibhuprasad Mahakud , Bibhuti Parida , Ipsit Panda , Souvik Maity , Arpita Sahoo , Reeta Sharma

We revisit benchmarks for differentially private image classification. We suggest a comprehensive set of benchmarks, allowing researchers to evaluate techniques for differentially private machine learning in a variety of settings, including…

Machine Learning · Computer Science 2026-01-28 Sabrina Mokhtari , Sara Kodeiri , Shubhankar Mohapatra , Florian Tramèr , Gautam Kamath

The development of Digital Twins (DTs) represents a transformative advance for simulating and optimizing complex systems in a controlled digital space. Despite their potential, the challenge of constructing DTs that accurately replicate and…

Systems and Control · Electrical Eng. & Systems 2024-06-21 Longfei Ma , Nan Cheng , Xiucheng Wang , Jiong Chen , Yinjun Gao , Dongxiao Zhang , Jun-Jie Zhang

Research exploring how to support decision-making has often used machine learning to automate or assist human decisions. We take an alternative approach for improving decision-making, using machine learning to help stakeholders surface ways…

Human-Computer Interaction · Computer Science 2023-02-24 Angie Zhang , Olympia Walker , Kaci Nguyen , Jiajun Dai , Anqing Chen , Min Kyung Lee

A `state of the art' model A surpasses humans in a benchmark B, but fails on similar benchmarks C, D, and E. What does B have that the other benchmarks do not? Recent research provides the answer: spurious bias. However, developing A to…

Computation and Language · Computer Science 2020-08-11 Swaroop Mishra , Anjana Arunkumar , Bhavdeep Sachdeva , Chris Bryan , Chitta Baral

Model-based experimental design is attracting increasing attention in chemical process engineering. Typically, an iterative procedure is pursued: an approximate model is devised, prescribed experiments are then performed and the resulting…

Optimization and Control · Mathematics 2021-01-25 Charlie Vanaret , Philipp Seufert , Jan Schwientek , Gleb Karpov , Gleb Ryzhakov , Ivan Oseledets , Norbert Asprion , Michael Bortz

We propose an approach for dynamic efficiency evaluation across multiple organizational dimensions using data envelopment analysis (DEA). The method generates both dimension-specific and aggregate efficiency scores, incorporates desirable…

Optimization and Control · Mathematics 2026-04-07 Hashem Omrani , Raha Imanirad , Adam Diamant , Utkarsh Verma , Amol Verma , Fahad Razak

Rigorous and reproducible evaluation is critical for assessing the state of the art and for guiding scientific advances in Artificial Intelligence. Evaluation is challenging in practice due to several reasons, including benchmark…

Current benchmarks that test LLMs on static, already-solved problems (e.g., math word problems) effectively demonstrated basic capability acquisition. The natural progression has been toward larger, more comprehensive and challenging…

Machine Learning · Computer Science 2025-12-15 Alwin Jin , Sean M. Hendryx , Vaskar Nath

We present a framework for designing scores to summarize performance metrics. Our design has two multi-criteria objectives: (1) improving on scores should improve all performance metrics, and (2) achieving pareto-optimal scores should…

Computers and Society · Computer Science 2024-10-10 Anmol Kabra , Mina Karzand , Tosca Lechner , Nathan Srebro , Serena Wang

Benchmark experiments are required to test, compare, tune, and understand optimization algorithms. Ideally, benchmark problems closely reflect real-world problem behavior. Yet, real-world problems are not always readily available for…

Neural and Evolutionary Computing · Computer Science 2020-08-17 Martin Zaefferer , Frederik Rehbach

I propose a new model to measure simultaneously academic and financial performances of scientific activities quantitatively. The tool is very simple and can be applied to any branches of science, while it is also adjustable to varying…

Physics and Society · Physics 2007-05-23 L. T. Handoko
‹ Prev 1 4 5 6 7 8 10 Next ›