English
Related papers

Related papers: Competitions in AI -- Robustly Ranking Solvers Usi…

200 papers

Although being a crucial question for the development of machine learning algorithms, there is still no consensus on how to compare classifiers over multiple data sets with respect to several criteria. Every comparison framework is…

Machine Learning · Statistics 2023-07-06 Christoph Jansen , Malte Nalenz , Georg Schollmeyer , Thomas Augustin

While games have been used extensively as milestones to evaluate game-playing AI, there exists no standardised framework for reporting the obtained observations. As a result, it remains difficult to draw general conclusions about the…

Artificial Intelligence · Computer Science 2020-07-07 Vanessa Volz , Boris Naujoks

Previous work on the competitive retrieval setting focused on a single-query setting: document authors manipulate their documents so as to improve their future ranking for a given query. We study a competitive setting where authors opt to…

Information Retrieval · Computer Science 2024-04-16 Haya Nachimovsky , Moshe Tennenholtz , Fiana Raiber , Oren Kurland

Imbalanced problems can arise in different real-world situations, and to address this, certain strategies in the form of resampling or balancing algorithms are proposed. This issue has largely been studied in the context of classification,…

Machine Learning · Computer Science 2025-07-17 Juscimara G. Avelino , George D. C. Cavalcanti , Rafael M. O. Cruz

For many queries in the Web retrieval setting there is an on-going ranking competition: authors manipulate their documents so as to promote them in rankings. Such competitions can have unwarranted effects not only in terms of retrieval…

Information Retrieval · Computer Science 2018-06-14 Gregory Goren , Oren Kurland , Moshe Tennenholtz , Fiana Raiber

Evaluating performance across optimization algorithms on many problems presents a complex challenge due to the diversity of numerical scales involved. Traditional data processing methods, such as hypothesis testing and Bayesian inference,…

Optimization and Control · Mathematics 2024-09-10 Yunpeng Jinng , Qunfeng Liu

Recent years have seen a significant surge in complex AI systems for competitive programming, capable of performing at admirable levels against human competitors. While steady progress has been made, the highest percentiles still remain out…

Machine Learning · Computer Science 2024-12-02 Petar Veličković , Alex Vitvitskyi , Larisa Markeeva , Borja Ibarz , Lars Buesing , Matej Balog , Alexander Novikov

In ranking competitions, document authors compete for the highest rankings by modifying their content in response to past rankings. Previous studies focused on human participants, primarily students, in controlled settings. The rise of…

Information Retrieval · Computer Science 2025-02-18 Tommy Mordo , Tomer Kordonsky , Haya Nachimovsky , Moshe Tennenholtz , Oren Kurland

In this position paper, we observe that empirical evaluation in Generative AI is at a crisis point since traditional ML evaluation and benchmarking strategies are insufficient to meet the needs of evaluating modern GenAI models and systems.…

Machine learning models play a key role for service providers looking to gain market share in consumer markets. However, traditional learning approaches do not take into account the existence of additional providers, who compete with each…

Machine Learning · Computer Science 2025-08-15 Ohad Einav , Nir Rosenfeld

Fair algorithm evaluation is conditioned on the existence of high-quality benchmark datasets that are non-redundant and are representative of typical optimization scenarios. In this paper, we evaluate three heuristics for selecting diverse…

Neural and Evolutionary Computing · Computer Science 2022-04-26 Gjorgjina Cenikj , Ryan Dieter Lang , Andries Petrus Engelbrecht , Carola Doerr , Peter Korošec , Tome Eftimov

International benchmarking competitions have become fundamental for the comparative performance assessment of image analysis methods. However, little attention has been given to investigating what can be learnt from these competitions. Do…

Computer Vision and Pattern Recognition · Computer Science 2023-04-03 Matthias Eisenmann , Annika Reinke , Vivienn Weru , Minu Dietlinde Tizabi , Fabian Isensee , Tim J. Adler , Sharib Ali , Vincent Andrearczyk , Marc Aubreville , Ujjwal Baid , Spyridon Bakas , Niranjan Balu , Sophia Bano , Jorge Bernal , Sebastian Bodenstedt , Alessandro Casella , Veronika Cheplygina , Marie Daum , Marleen de Bruijne , Adrien Depeursinge , Reuben Dorent , Jan Egger , David G. Ellis , Sandy Engelhardt , Melanie Ganz , Noha Ghatwary , Gabriel Girard , Patrick Godau , Anubha Gupta , Lasse Hansen , Kanako Harada , Mattias Heinrich , Nicholas Heller , Alessa Hering , Arnaud Huaulmé , Pierre Jannin , Ali Emre Kavur , Oldřich Kodym , Michal Kozubek , Jianning Li , Hongwei Li , Jun Ma , Carlos Martín-Isla , Bjoern Menze , Alison Noble , Valentin Oreiller , Nicolas Padoy , Sarthak Pati , Kelly Payette , Tim Rädsch , Jonathan Rafael-Patiño , Vivek Singh Bawa , Stefanie Speidel , Carole H. Sudre , Kimberlin van Wijnen , Martin Wagner , Donglai Wei , Amine Yamlahi , Moi Hoon Yap , Chun Yuan , Maximilian Zenk , Aneeq Zia , David Zimmerer , Dogu Baran Aydogan , Binod Bhattarai , Louise Bloch , Raphael Brüngel , Jihoon Cho , Chanyeol Choi , Qi Dou , Ivan Ezhov , Christoph M. Friedrich , Clifton Fuller , Rebati Raman Gaire , Adrian Galdran , Álvaro García Faura , Maria Grammatikopoulou , SeulGi Hong , Mostafa Jahanifar , Ikbeom Jang , Abdolrahim Kadkhodamohammadi , Inha Kang , Florian Kofler , Satoshi Kondo , Hugo Kuijf , Mingxing Li , Minh Huan Luu , Tomaž Martinčič , Pedro Morais , Mohamed A. Naser , Bruno Oliveira , David Owen , Subeen Pang , Jinah Park , Sung-Hong Park , Szymon Płotka , Elodie Puybareau , Nasir Rajpoot , Kanghyun Ryu , Numan Saeed , Adam Shephard , Pengcheng Shi , Dejan Štepec , Ronast Subedi , Guillaume Tochon , Helena R. Torres , Helene Urien , João L. Vilaça , Kareem Abdul Wahid , Haojie Wang , Jiacheng Wang , Liansheng Wang , Xiyue Wang , Benedikt Wiestler , Marek Wodzinski , Fangfang Xia , Juanying Xie , Zhiwei Xiong , Sen Yang , Yanwu Yang , Zixuan Zhao , Klaus Maier-Hein , Paul F. Jäger , Annette Kopp-Schneider , Lena Maier-Hein

Reinforcement learning competitions have formed the basis for standard research benchmarks, galvanized advances in the state-of-the-art, and shaped the direction of the field. Despite this, a majority of challenges suffer from the same…

Patterns of wins and losses in pairwise contests, such as occur in sports and games, consumer research and paired comparison studies, and human and animal social hierarchies, are commonly analyzed using probabilistic models that allow one…

Physics and Society · Physics 2025-11-03 Maximilian Jerdee , M. E. J. Newman

The stochastic nature of iterative optimization heuristics leads to inherently noisy performance measurements. Since these measurements are often gathered once and then used repeatedly, the number of collected samples will have a…

Neural and Evolutionary Computing · Computer Science 2022-04-25 Diederick Vermetten , Hao Wang , Manuel López-Ibañez , Carola Doerr , Thomas Bäck

Public AI benchmark results are widely broadcast by model developers as indicators of model quality within a growing and competitive market. However, these advertised scores do not necessarily reflect the traits of interest to those who…

Artificial intelligent (AI) algorithms, such as deep learning and XGboost, are used in numerous applications including computer vision, autonomous driving, and medical diagnostics. The robustness of these AI algorithms is of great interest…

Machine Learning · Statistics 2020-10-30 Jiayi Lian , Laura Freeman , Yili Hong , Xinwei Deng

Artificial Intelligence (AI) is a fast-growing research and development (R&D) discipline which is attracting increasing attention because of its promises to bring vast benefits for consumers and businesses, with considerable benefits…

Artificial Intelligence · Computer Science 2022-05-10 Zhenghua Chen , Min Wu , Alvin Chan , Xiaoli Li , Yew-Soon Ong

The small sample imbalance (S&I) problem is a major challenge in machine learning and data analysis. It is characterized by a small number of samples and an imbalanced class distribution, which leads to poor model performance. In addition,…

Machine Learning · Computer Science 2025-04-22 Shuxian Zhao , Jie Gui , Minjing Dong , Baosheng Yu , Zhipeng Gui , Lu Dong , Yuan Yan Tang , James Tin-Yau Kwok

Classification algorithms based on Artificial Intelligence (AI) are nowadays applied in high-stakes decisions in finance, healthcare, criminal justice, or education. Individuals can strategically adapt to the information gathered about…

Computer Science and Game Theory · Computer Science 2025-08-14 Marta C. Couto , Flavia Barsotti , Fernando P. Santos