English
Related papers

Related papers: The Dataset Multiplicity Problem: How Unreliable D…

200 papers

Counterfactual explanations are widely used to interpret machine learning predictions by identifying minimal changes to input features that would alter a model's decision. However, most existing counterfactual methods have not been tested…

Machine Learning · Computer Science 2026-02-03 Leonidas Christodoulou , Chang Sun

Language models have shown promise in various tasks but can be affected by undesired data during training, fine-tuning, or alignment. For example, if some unsafe conversations are wrongly annotated as safe ones, the model fine-tuned on…

Machine Learning · Computer Science 2024-03-26 Zhaowei Zhu , Jialu Wang , Hao Cheng , Yang Liu

Data heterogeneity plays a pivotal role in determining the performance of machine learning (ML) systems. Traditional algorithms, which are typically designed to optimize average performance, often overlook the intrinsic diversity within…

Machine Learning · Computer Science 2025-06-03 Jiashuo Liu , Peng Cui

We investigate the problem of machine learning with mislabeled training data. We try to make the effects of mislabeled training better understood through analysis of the basic model and equations that characterize the problem. This includes…

Machine Learning · Computer Science 2019-09-23 Herbert Gish , Jan Silovsky , Man-Ling Sung , Man-Hung Siu , William Hartmann , Zhuolin Jiang

Large Language Models (LLMs) annotated datasets are widely used nowadays, however, large-scale annotations often show biases in low-quality datasets. For example, Multiple-Choice Questions (MCQs) datasets with one single correct option is…

Computation and Language · Computer Science 2026-01-08 Zipeng Ling , Shuliang Liu , Yuehao Tang , Chen Huang , Gaoyang Jiang , Shenghong Fu , Junqi Yang , Yao Wan , Jiawan Zhang , Kejia Huang , Xuming Hu

Many data mining approaches aim at modelling and predicting human behaviour. An important quantity of interest is the quality of model-based predictions, e.g. for finding a competition winner with best prediction performance. In real life,…

Human-Computer Interaction · Computer Science 2017-02-27 Kevin Jasberg , Sergej Sizov

Model multiplicity, the phenomenon where multiple models achieve similar performance despite different underlying learned functions, introduces arbitrariness in model selection. While this arbitrariness may seem inconsequential in…

Computers and Society · Computer Science 2024-09-16 Prakhar Ganesh , Ihsan Ibrahim Daldaban , Ignacio Cofone , Golnoosh Farnadi

The performance of machine learning models relies heavily on the quality of input data, yet real-world applications often face significant data-related challenges. A common issue arises when curating training data or deploying models: two…

Machine Learning · Computer Science 2025-09-24 Varun Babbar , Zhicheng Guo , Cynthia Rudin

This paper addresses a multi-label predictive fault classification problem for multidimensional time-series data. While fault (event) detection problems have been thoroughly studied in literature, most of the state-of-the-art techniques…

Machine Learning · Computer Science 2020-01-29 Wenyu Zhang , Devesh K. Jha , Emil Laftchiev , Daniel Nikovski

While data-driven predictive models are a strictly technological construct, they may operate within a social context in which benign engineering choices entail implicit, indirect and unexpected real-life consequences. Fairness of such…

Machine Learning · Computer Science 2024-07-11 Kacper Sokol , Meelis Kull , Jeffrey Chan , Flora Salim

In this study we provide empirical evidence demonstrating that the quality of training data impacts model performance in Human Pose Estimation (HPE). Inaccurate labels in widely used data sets, ranging from minor errors to severe…

Computer Vision and Pattern Recognition · Computer Science 2025-09-25 Arnold Schwarz , Levente Hernadi , Felix Bießmann , Kristian Hildebrand

Segmentation uncertainty models predict a distribution over plausible segmentations for a given input, which they learn from the annotator variation in the training set. However, in practice these annotations can differ systematically in…

Computer Vision and Pattern Recognition · Computer Science 2023-03-29 Kilian Zepf , Eike Petersen , Jes Frellsen , Aasa Feragen

Data-driven algorithms play a large role in decision making across a variety of industries. Increasingly, these algorithms are being used to make decisions that have significant ramifications for people's social and economic well-being,…

Machine Learning · Computer Science 2018-09-26 J. Henry Hinnefeld , Peter Cooman , Nat Mammo , Rupert Deese

Deep Neural Networks are well known for efficiently fitting training data, yet experiencing poor generalization capabilities whenever some kind of bias dominates over the actual task labels, resulting in models learning "shortcuts". In…

Machine Learning · Computer Science 2024-08-12 Pietro Morerio , Ruggero Ragonesi , Vittorio Murino

Text style transfer is an exciting task within the field of natural language generation that is often plagued by the need for high-quality paired datasets. Furthermore, training a model for multi-attribute text style transfer requires…

Computation and Language · Computer Science 2023-05-26 Debarati Das , David Ma , Dongyeop Kang

Artificial Intelligence (AI) systems are not intrinsically neutral and biases trickle in any type of technological tool. In particular when dealing with people, the impact of AI algorithms' technical errors originating with mislabeled data…

Artificial Intelligence · Computer Science 2025-04-03 Camilla Quaresmini , Giuseppe Primiero

Stochastic simulation is widely used to study complex systems composed of various interconnected subprocesses, such as input processes, routing and control logic, optimization routines, and data-driven decision modules. In practice, these…

Computation · Statistics 2026-02-19 Mohammadmahdi Ghasemloo , David J. Eckman , Yaxian Li

Machine learning (ML) systems are increasingly deployed in high-stakes domains where reliability is paramount. This thesis investigates how uncertainty estimation can enhance the safety and trustworthiness of ML, focusing on selective…

Machine Learning · Computer Science 2025-09-09 Stephan Rabanser

Labeled datasets reflect the biases of their annotation pipelines, which sometimes introduce label bias: group-conditional label errors that cause systematic performance disparities across demographic subgroups. Label bias in image…

Computer Vision and Pattern Recognition · Computer Science 2026-05-11 Aditya Parikh , Stella Frank , Sneha Das , Aasa Feragen

The adaptation and use of Machine Learning (ML) in our daily lives has led to concerns in lack of transparency, privacy, reliability, among others. As a result, we are seeing research in niche areas such as interpretability, causality, bias…

Machine Learning · Computer Science 2024-06-04 Fahimeh Fakour , Ali Mosleh , Ramin Ramezani
‹ Prev 1 3 4 5 6 7 10 Next ›