中文
相关论文

相关论文: FIB: A Method for Evaluation of Feature Impact Bal…

200 篇论文

We encounter variables with little variation often in educational data mining (EDM) due to the demographics of higher education and the questions we ask. Yet, little work has examined how to analyze such data. Therefore, we conducted a…

统计方法学 · 统计学 2022-01-12 Nicholas T. Young , Marcos D. Caballero

A/B testing is gaining attention in the automotive sector as a promising tool to measure causal effects from software changes. Different from the web-facing businesses, where A/B testing has been well-established, the automotive domain…

软件工程 · 计算机科学 2021-11-12 Yuchu Liu , David Issa Mattos , Jan Bosch , Helena Holmström Olsson , Jonn Lantz

We characterise the unbiasedness of the score function, viewed as an inference function for a class of finite mixture models. The models studied represent the situation where there is a stratification of the observations in a finite number…

统计理论 · 数学 2023-05-16 Rodrigo Labouriau

The aim of Active Learning is to select the most informative samples from an unlabelled set of data. This is useful in cases where the amount of data is large and labelling is expensive, such as in machine vision or medical imaging. Two…

计算机视觉与模式识别 · 计算机科学 2026-01-13 Julien Combes , Alexandre Derville , Jean-François Coeurjolly

This paper shows that two commonly used evaluation metrics for generative models, the Fr\'echet Inception Distance (FID) and the Inception Score (IS), are biased -- the expected value of the score computed for a finite sample set is not the…

计算机视觉与模式识别 · 计算机科学 2020-06-17 Min Jin Chong , David Forsyth

Although many fairness criteria have been proposed to ensure that machine learning algorithms do not exhibit or amplify our existing social biases, these algorithms are trained on datasets that can themselves be statistically biased. In…

机器学习 · 计算机科学 2023-05-04 Yiqiao Liao , Parinaz Naghizadeh

This study evaluates fine-tuning strategies for text classification using the DistilBERT model, specifically the distilbert-base-uncased-finetuned-sst-2-english variant. Through structured experiments, we examine the influence of…

计算与语言 · 计算机科学 2025-01-03 Giuliano Lorenzoni , Ivens Portugal , Paulo Alencar , Donald Cowan

It is widely held that one cause of downstream bias in classifiers is bias present in the training data. Rectifying such biases may involve context-dependent interventions such as training separate models on subgroups, removing features…

机器学习 · 计算机科学 2024-06-04 Peter W. Chang , Leor Fishman , Seth Neel

Record matching models typically output a real-valued matching score that is later consumed through thresholding, ranking, or human review. While fairness in record matching has mostly been assessed using binary decisions at a fixed…

机器学习 · 计算机科学 2026-02-24 Mohammad Hossein Moslemi , Mostafa Milani

The quality and generality of deep image features is crucially determined by the data they have been trained on, but little is known about this often overlooked effect. In this paper, we systematically study the effect of variations in the…

计算机视觉与模式识别 · 计算机科学 2020-07-20 Othman Sbai , Camille Couprie , Mathieu Aubry

Nowadays, feature selection is frequently used in machine learning when there is a risk of performance degradation due to overfitting or when computational resources are limited. During the feature selection process, the subset of features…

机器学习 · 计算机科学 2023-01-02 Sergey A. Saltykov

The class-imbalance issue is intrinsic to many real-world machine learning tasks, particularly to the rare-event classification problems. Although the impact and treatment of imbalanced data is widely known, the magnitude of a metric's…

机器学习 · 计算机科学 2022-06-22 Azim Ahmadzadeh , Rafal A. Angryk

Context: Software engineering researchers have undertaken many experiments investigating the potential of software defect prediction algorithms. Unfortunately, some widely used performance metrics are known to be problematic, most notably…

软件工程 · 计算机科学 2021-06-23 Jingxiu Yao , Martin Shepperd

Influence functions (IFs) are a powerful tool for detecting anomalous examples in large scale datasets. However, they are unstable when applied to deep networks. In this paper, we provide an explanation for the instability of IFs and…

Model bias triggered by long-tailed data has been widely studied. However, measure based on the number of samples cannot explicate three phenomena simultaneously: (1) Given enough data, the classification performance gain is marginal with…

计算机视觉与模式识别 · 计算机科学 2023-04-11 Yanbiao Ma , Licheng Jiao , Fang Liu , Yuxin Li , Shuyuan Yang , Xu Liu

This paper introduces a novel perspective about error in machine learning and proposes inverse feature learning (IFL) as a representation learning approach that learns a set of high-level features based on the representation of error for…

机器学习 · 计算机科学 2020-03-10 Behzad Ghazanfari , Fatemeh Afghah

Resilience to class imbalance and confounding biases, together with the assurance of fairness guarantees are highly desirable properties of autonomous decision-making systems with real-life impact. Many different targeted solutions have…

机器学习 · 计算机科学 2021-05-14 Elisa Ferrari , Davide Bacciu

In a binary classification problem the feature vector (predictor) is the input to a scoring function that produces a decision value (score), which is compared to a particular chosen threshold to provide a final class prediction (output).…

机器学习 · 计算机科学 2021-11-11 Waleed A. Yousef

Data imbalance is a fundamental challenge in applying language models to biomedical applications, particularly in ICD code prediction tasks where label and demographic distributions are uneven. While state-of-the-art language models have…

机器学习 · 计算机科学 2025-02-17 Precious Jones , Weisi Liu , I-Chan Huang , Xiaolei Huang

The property of conformal predictors to guarantee the required accuracy rate makes this framework attractive in various practical applications. However, this property is achieved at a price of reduction in precision. In the case of…

机器学习 · 计算机科学 2021-08-13 Marharyta Aleksandrova , Oleg Chertov