English
Related papers

Related papers: $F_\beta$-plot -- a visual tool for evaluating imb…

200 papers

As the social impact of visual recognition has been under scrutiny, several protected-attribute balanced datasets emerged to address dataset bias in imbalanced datasets. However, in facial attribute classification, dataset bias stems from…

Computer Vision and Pattern Recognition · Computer Science 2022-09-16 Jiazhi Li , Wael Abd-Almageed

Model averaging is a useful and robust method for dealing with model uncertainty in statistical analysis. Often, it is useful to consider data subset selection at the same time, in which model selection criteria are used to compare models…

Methodology · Statistics 2023-10-26 Ethan T. Neil , Jacob W. Sitison

The multiple-biomarker classifier problem and its assessment are reviewed against the background of some fundamental principles from the field of statistical pattern recognition, machine learning, or the recently so-called "data science". A…

Genomics · Quantitative Biology 2019-11-01 Waleed A. Yousef

The purpose of this project was to collect and analyse data about the comparability and real-life applicability of published results focusing on Microsoft Windows malware, more specifically the impact of dataset size and testing dataset…

Cryptography and Security · Computer Science 2022-06-14 David Illes

Multi-label classification is a widely encountered problem in daily life, where an instance can be associated with multiple classes. In theory, this is a supervised learning method that requires a large amount of labeling. However,…

Computer Vision and Pattern Recognition · Computer Science 2023-08-02 XIn Zhang , Yuqi Song , Fei Zuo , Xiaofeng Wang

In cases of uncertainty, a multi-class classifier preferably returns a set of candidate classes instead of predicting a single class label with little guarantee. More precisely, the classifier should strive for an optimal balance between…

Machine Learning · Computer Science 2020-05-28 Thomas Mortier , Marek Wydmuch , Krzysztof Dembczyński , Eyke Hüllermeier , Willem Waegeman

Resilience to class imbalance and confounding biases, together with the assurance of fairness guarantees are highly desirable properties of autonomous decision-making systems with real-life impact. Many different targeted solutions have…

Machine Learning · Computer Science 2021-05-14 Elisa Ferrari , Davide Bacciu

In applications with significant class imbalance or asymmetric costs, metrics such as the $F_\beta$-measure, AM measure, Jaccard similarity coefficient, and weighted accuracy offer more suitable evaluation criteria than standard binary…

Machine Learning · Computer Science 2025-12-30 Anqi Mao , Mehryar Mohri , Yutao Zhong

Deep Neural networks have gained lots of attention in recent years thanks to the breakthroughs obtained in the field of Computer Vision. However, despite their popularity, it has been shown that they provide limited robustness in their…

Machine Learning · Computer Science 2020-08-21 Marco Maggipinto , Matteo Terzi , Gian Antonio Susto

Data imbalance is a well-known issue in the field of machine learning, attributable to the cost of data collection, the difficulty of labeling, and the geographical distribution of the data. In computer vision, bias in data distribution…

Computer Vision and Pattern Recognition · Computer Science 2023-08-23 Shubham Shrivastava , Xianling Zhang , Sushruth Nagesh , Armin Parchami

Surveys are commonly used to facilitate research in epidemiology, health, and the social and behavioral sciences. Often, these surveys are not simple random samples, and respondents are given weights reflecting their probability of…

Methodology · Statistics 2024-08-20 Adway S. Wadekar , Jerome P. Reiter

In visual generation tasks, the responses and combinations of complex concepts often lack stability and are error-prone, which remains an under-explored area. In this paper, we attempt to explore the causal factors for poor concept…

Computer Vision and Pattern Recognition · Computer Science 2025-11-12 Yukai Shi , Jiarong Ou , Rui Chen , Haotian Yang , Jiahao Wang , Xin Tao , Pengfei Wan , Di Zhang , Kun Gai

In many classification settings, the class of primary interest is underrepresented, leading to imbalanced data problems that arise in applications such as rare disease detection and fraud identification. In these contexts, identifying a…

Machine Learning · Statistics 2026-05-06 Daniel Fraiman , Ricardo Fraiman

This paper addresses a multi-label predictive fault classification problem for multidimensional time-series data. While fault (event) detection problems have been thoroughly studied in literature, most of the state-of-the-art techniques…

Machine Learning · Computer Science 2020-01-29 Wenyu Zhang , Devesh K. Jha , Emil Laftchiev , Daniel Nikovski

Random-effects meta-analyses of observational studies can produce biased estimates if the synthesized studies are subject to unmeasured confounding. We propose sensitivity analyses quantifying the extent to which unmeasured confounding of…

Methodology · Statistics 2017-10-10 Maya B. Mathur , Tyler J. VanderWeele

We present a new approach, called meta-meta classification, to learning in small-data settings. In this approach, one uses a large set of learning problems to design an ensemble of learners, where each learner has high bias and low variance…

Machine Learning · Computer Science 2020-06-16 Arkabandhu Chowdhury , Dipak Chaudhari , Swarat Chaudhuri , Chris Jermaine

While the importance of automatic image analysis is continuously increasing, recent meta-research revealed major flaws with respect to algorithm validation. Performance metrics are particularly key for meaningful, objective, and transparent…

Image and Video Processing · Electrical Eng. & Systems 2023-12-08 Annika Reinke , Minu D. Tizabi , Carole H. Sudre , Matthias Eisenmann , Tim Rädsch , Michael Baumgartner , Laura Acion , Michela Antonelli , Tal Arbel , Spyridon Bakas , Peter Bankhead , Arriel Benis , Matthew Blaschko , Florian Buettner , M. Jorge Cardoso , Jianxu Chen , Veronika Cheplygina , Evangelia Christodoulou , Beth Cimini , Gary S. Collins , Sandy Engelhardt , Keyvan Farahani , Luciana Ferrer , Adrian Galdran , Bram van Ginneken , Ben Glocker , Patrick Godau , Robert Haase , Fred Hamprecht , Daniel A. Hashimoto , Doreen Heckmann-Nötzel , Peter Hirsch , Michael M. Hoffman , Merel Huisman , Fabian Isensee , Pierre Jannin , Charles E. Kahn , Dagmar Kainmueller , Bernhard Kainz , Alexandros Karargyris , Alan Karthikesalingam , A. Emre Kavur , Hannes Kenngott , Jens Kleesiek , Andreas Kleppe , Sven Kohler , Florian Kofler , Annette Kopp-Schneider , Thijs Kooi , Michal Kozubek , Anna Kreshuk , Tahsin Kurc , Bennett A. Landman , Geert Litjens , Amin Madani , Klaus Maier-Hein , Anne L. Martel , Peter Mattson , Erik Meijering , Bjoern Menze , David Moher , Karel G. M. Moons , Henning Müller , Brennan Nichyporuk , Felix Nickel , M. Alican Noyan , Jens Petersen , Gorkem Polat , Susanne M. Rafelski , Nasir Rajpoot , Mauricio Reyes , Nicola Rieke , Michael Riegler , Hassan Rivaz , Julio Saez-Rodriguez , Clara I. Sánchez , Julien Schroeter , Anindo Saha , M. Alper Selver , Lalith Sharan , Shravya Shetty , Maarten van Smeden , Bram Stieltjes , Ronald M. Summers , Abdel A. Taha , Aleksei Tiulpin , Sotirios A. Tsaftaris , Ben Van Calster , Gaël Varoquaux , Manuel Wiesenfarth , Ziv R. Yaniv , Paul Jäger , Lena Maier-Hein

With the rapid increase of large-scale, real-world datasets, it becomes critical to address the problem of long-tailed data distribution (i.e., a few classes account for most of the data, while most classes are under-represented). Existing…

Computer Vision and Pattern Recognition · Computer Science 2019-01-18 Yin Cui , Menglin Jia , Tsung-Yi Lin , Yang Song , Serge Belongie

We show that established performance metrics in binary classification, such as the F-score, the Jaccard similarity coefficient or Matthews' correlation coefficient (MCC), are not robust to class imbalance in the sense that if the proportion…

Machine Learning · Statistics 2024-04-12 Hajo Holzmann , Bernhard Klar

Classification tasks in machine learning involving more than two classes are known by the name of "multi-class classification". Performance indicators are very useful when the aim is to evaluate and compare different classification models…

Machine Learning · Statistics 2020-08-14 Margherita Grandini , Enrico Bagli , Giorgio Visani