English
Related papers

Related papers: The impact of using biased performance metrics on …

200 papers

Several important aspects of software product quality can be evaluated using dynamic metrics that effectively capture and reflect the software's true runtime behavior. While the extent of research in this field is still relatively limited,…

Software Engineering · Computer Science 2021-01-12 Amjed Tahir , Stephen G. MacDonell

Causal graphs are widely used in software engineering to document and explore causal relationships. Though widely used, they may also be wildly misleading. Causal structures generated from SE data can be highly variable. This instability is…

Software Engineering · Computer Science 2025-05-20 Jeremy Hulse , Nasir U. Eisty , Tim Menzies

Prevalent Fault Localization (FL) techniques rely on tests to localize buggy program elements. Tests could be treated as fuel to further boost FL by providing more debugging information. Therefore, it is highly valuable to measure the Fault…

Software Engineering · Computer Science 2025-01-07 Yifan Zhao , Zeyu Sun , Guoqing Wang , Qingyuan Liang , Yakun Zhang , Yiling Lou , Dan Hao , Lu Zhang

In defect prediction community, many defect prediction models have been proposed and indeed more new models are continuously being developed. However, there is no consensus on how to evaluate the performance of a newly proposed model. In…

Software Engineering · Computer Science 2023-02-07 Xutong Liu , Shiran Liu , Zhaoqiang Guo , Peng Zhag , Yibiao Yang , Huihui Liu , Hongmin Lu , Yanhui Li , Lin Chen , Yuming Zhou

One source of software project challenges and failures is the systematic errors introduced by human cognitive biases. Although extensively explored in cognitive psychology, investigations concerning cognitive biases have only recently…

Software Engineering · Computer Science 2022-03-22 Rahul Mohanani , Iflaah Salman , Burak Turhan , Pilar Rodriguez , Paul Ralph

Empirical and LLM-based research in model-driven engineering increasingly relies on datasets of software models, for instance, to train or evaluate machine learning techniques for modeling support. These datasets have a significant impact…

Software Engineering · Computer Science 2026-03-06 Philipp-Lorenz Glaser , Lola Burgueño , Dominik Bork

Detecting and mitigating bias in speaker verification systems is important, as datasets, processing choices and algorithms can lead to performance differences that systematically favour some groups of people while disadvantaging others.…

Audio and Speech Processing · Electrical Eng. & Systems 2024-08-27 Wiebke Hutiri , Tanvina Patel , Aaron Yi Ding , Odette Scharenborg

The estimation of the amount of uncertainty featured by predictive machine learning models has acquired a great momentum in recent years. Uncertainty estimation provides the user with augmented information about the model's confidence in…

Machine Learning · Computer Science 2022-10-31 Ibai Laña , Ignacio , Olabarrieta , Javier Del Ser

Cross-project defect prediction (CPDP) has been deemed as an emerging technology of software quality assurance, especially in new or inactive projects, and a few improved methods have been proposed to support better defect prediction.…

Software Engineering · Computer Science 2014-11-18 Peng He , Bing Li , Yutao Ma

To face future reliability challenges, it is necessary to quantify the risk of error in any part of a computing system. To this goal, the Architectural Vulnerability Factor (AVF) has long been used for chips. However, this metric is used…

Hardware Architecture · Computer Science 2023-08-02 Luc Jaulmes , Miquel Moretó , Mateo Valero , Marc Casas

When a model's performance differs across socially or culturally relevant groups--like race, gender, or the intersections of many such groups--it is often called "biased." While much of the work in algorithmic fairness over the last several…

Methodology · Statistics 2022-07-01 Kristian Lum , Yunfeng Zhang , Amanda Bower

Code metrics are easy to define, but not so easy to justify. It is hard to prove that a metric is valid, i.e., that measured numerical values imply anything on the vaguely defined, yet crucial software properties such as complexity and…

Software Engineering · Computer Science 2012-01-17 Joseph Gil , Maayan Goldstein , Dany Moshkovich

High-dimensional data acquired from biological experiments such as next generation sequencing are subject to a number of confounding effects. These effects include both technical effects, such as variation across batches from instrument…

Machine Learning · Computer Science 2018-12-11 Kabir Manghnani , Adam Drake , Nathan Wan , Imran Haque

Dynamical systems are frequently used to model biological systems. When these models are fit to data it is necessary to ascertain the uncertainty in the model fit. Here we present prediction deviation, a new metric of uncertainty that…

Applications · Statistics 2017-06-08 Benjamin Letham , Portia A. Letham , Cynthia Rudin , Edward P. Browne

Applying software defect esimation techniques and presenting this information in a compact and impactful decision table can clearly illustrate to collaborative groups how critical this position is in the overall development cycle. The Test…

Software Engineering · Computer Science 2007-11-13 James Cusick

Context: Large language models (LLMs) are increasingly used to screen literature for systematic reviews (SRs), but the standard confusion-matrix metrics used to evaluate them can mislead under the imbalanced, cost-asymmetric conditions of…

Software Engineering · Computer Science 2026-04-28 Lech Madeyski , Barbara Kitchenham , Martin Shepperd

Quantitative research relies heavily on coding, and coding errors are relatively common even in published research. In this paper, we examine whether individuals are more or less likely to check their code depending on the results they…

General Economics · Economics 2025-09-26 Bruno Ferman , Lucas Finamor

Quantitative metrics derived from software repositories and package ecosystems are widely used to assess the impact, popularity, maintenance, and criticality of free and open source software (FOSS) projects. However, these metrics are often…

Cryptography and Security · Computer Science 2026-05-21 Ben Swierzy , Timo Pohl , Marc Ohm , Michael Meier

Clinical machine learning applications are often plagued with confounders that are clinically irrelevant, but can still artificially boost the predictive performance of the algorithms. Confounding is especially problematic in mobile health…

Applications · Statistics 2018-11-29 Elias Chaibub Neto

Prediction algorithms that quantify the expected benefit of a given treatment conditional on patient characteristics can critically inform medical decisions. Quantifying the performance of treatment benefit prediction algorithms is an…

Methodology · Statistics 2023-05-17 Yuan Xia , Paul Gustafson , Mohsen Sadatsafavi