English
Related papers

Related papers: A Rosetta Stone for AI Benchmarks

200 papers

Artificial Intelligence (AI) has achieved significant advancements in technology and research with the development over several decades, and is widely used in many areas including computing vision, natural language processing, time-series…

AI evaluations have become critical tools for assessing large language model capabilities and safety. This paper presents practical insights from eight months of maintaining $inspect\_evals$, an open-source repository of 70+…

Computation and Language · Computer Science 2025-07-10 Alexandra Abbas , Celia Waggoner , Justin Olive

Dynamic benchmarks interweave model fitting and data collection in an attempt to mitigate the limitations of static benchmarks. In contrast to an extensive theoretical and empirical study of the static setting, the dynamic counterpart lags…

Machine Learning · Computer Science 2023-03-03 Ali Shirali , Rediet Abebe , Moritz Hardt

Benchmarking the performance of quantum optimization algorithms is crucial for identifying utility for industry-relevant use cases. Benchmarking processes vary between optimization applications and depend on user-specified goals. The…

AI researchers employ not only the scientific method, but also methodology from mathematics and engineering. However, the use of the scientific method - specifically hypothesis testing - in AI is typically conducted in service of…

The evaluation of academic theses is a cornerstone of higher education, ensuring rigor and integrity. Traditional methods, though effective, are time-consuming and subject to evaluator variability. This paper presents RubiSCoT, an…

Artificial Intelligence · Computer Science 2025-11-24 Thorsten Fröhlich , Tim Schlippe

Comprehensive benchmarking of clustering algorithms is rendered difficult by two key factors: (i)~the elusiveness of a unique mathematical definition of this unsupervised learning approach and (ii)~dependencies between the generating models…

Neural and Evolutionary Computing · Computer Science 2022-01-11 Cameron Shand , Richard Allmendinger , Julia Handl , Andrew Webb , John Keane

The automation of AI R&D (AIRDA) could have significant implications, but its extent and ultimate effects remain uncertain. We need empirical data to resolve these uncertainties, but existing data (primarily capability benchmarks) may not…

Computers and Society · Computer Science 2026-03-09 Alan Chan , Ranay Padarath , Joe Kwon , Hilary Greaves , Markus Anderljung

Conversational Artificial Intelligence (AI) systems have recently sky-rocketed in popularity and are now used in many applications, from car assistants to customer support. The development of conversational AI systems is supported by a…

Human-Computer Interaction · Computer Science 2020-12-23 Johan Aronsson , Philip Lu , Daniel Strüber , Thorsten Berger

Experimental evaluation is an integral part in the design process of algorithms. Publicly available benchmark instances are widely used to evaluate methods in SAT solving. For the interpretation of results and the design of algorithm…

Artificial Intelligence · Computer Science 2021-09-10 Markus Iser , Luca Springer , Carsten Sinz

Due to increasing amounts of data and compute resources, deep learning achieves many successes in various domains. The application of deep learning on the mobile and embedded devices is taken more and more attentions, benchmarking and…

Machine Learning · Computer Science 2020-05-12 Chunjie Luo , Xiwen He , Jianfeng Zhan , Lei Wang , Wanling Gao , Jiahui Dai

Algorithm evaluation and comparison are fundamental questions in machine learning and statistics -- how well does an algorithm perform at a given modeling task, and which algorithm performs best? Many methods have been developed to assess…

Statistics Theory · Mathematics 2025-11-25 Yuetian Luo , Rina Foygel Barber

To build general-purpose artificial intelligence systems that can deal with unknown variables across unknown domains, we need benchmarks that measure how well these systems perform on tasks they have never seen before. A prerequisite for…

Artificial Intelligence · Computer Science 2022-05-25 Gautham Venkatasubramanian , Sibesh Kar , Abhimanyu Singh , Shubham Mishra , Dushyant Yadav , Shreyansh Chandak

In the field of scientific computing, one often finds several alternative software packages (with open or closed source code) for solving a specific problem. These packages sometimes even use alternative methodological approaches, e.g.,…

In the future, most companies will be confronted with the topic of Artificial Intelligence (AI) and will have to decide on their strategy in this regards. Currently, a lot of companies are thinking about whether and how AI and the usage of…

Other Computer Science · Computer Science 2022-06-03 Jens Heidrich , Andreas Jedlitschka , Adam Trendowicz , Anna Maria Vollmer

Frontier AI safety policies highlight automation of AI research and development (R&D) by AI agents as an important capability to anticipate. However, there exist few evaluations for AI R&D capabilities, and none that are highly realistic…

Artificial intelligence (AI) models are increasingly autonomous in decision-making, making pursuing responsible AI more critical than ever. Responsible AI (RAI) is defined by its commitment to transparency, privacy, safety, inclusiveness,…

Computers and Society · Computer Science 2025-01-28 Gemma Galdon Clavell , Rubén González-Sendino , Paola Vazquez

Although artificial intelligence (AI) systems are becoming increasingly indispensable, research into how humans rely on these systems (AI reliance) is lagging behind. To advance this research, this survey presents a novel, comprehensive…

Human-Computer Interaction · Computer Science 2025-09-03 Sven Eckhardt , Niklas Kühl , Mateusz Dolata , Gerhard Schwabe

Artificial intelligence is more ubiquitous in multiple domains. Smartphones, social media platforms, search engines, and autonomous vehicles are just a few examples of applications that utilize artificial intelligence technologies to…

Machine Learning · Computer Science 2025-06-24 Teemu Niskanen , Tuomo Sipola , Olli Väänänen

It is common to evaluate the performance of a machine learning model by measuring its predictive power on a test dataset. This approach favors complicated models that can smoothly fit complex functions and generalize well from training data…

Machine Learning · Computer Science 2022-10-07 Hugo Cisneros , Josef Sivic , Tomas Mikolov