English
Related papers

Related papers: Position: Measure Dataset Diversity, Don't Just Cl…

200 papers

Research in machine learning (ML) has primarily argued that models trained on incomplete or biased datasets can lead to discriminatory outputs. In this commentary, we propose moving the research focus beyond bias-oriented framings by…

Human-Computer Interaction · Computer Science 2021-09-17 Milagros Miceli , Julian Posada , Tianling Yang

We identify the task of measuring data to quantitatively characterize the composition of machine learning data and datasets. Similar to an object's height, width, and volume, data measurements quantify different attributes of data along…

Currently, data and model size dominate the narrative in the training of super-large, powerful models. However, there has been a lack of exploration on the effect of other attributes of the training dataset on model performance. We…

Machine Learning · Computer Science 2025-01-22 Kavita Selva , Satita Vittayaareekul , Brando Miranda

There has been a surge of recent interest in sociocultural diversity in machine learning (ML) research, with researchers (i) examining the benefits of diversity as an organizational solution for alleviating problems with algorithmic bias,…

Computers and Society · Computer Science 2021-07-21 Sina Fazelpour , Maria De-Arteaga

The diversity of training datasets is usually perceived as an important aspect to obtain a robust model. However, the definition of diversity is often not defined or differs across papers, and while some metrics exist, the quantification of…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Théo Sourget , Niclas Claßen , Jack Junchi Xu , Rob van der Goot , Veronika Cheplygina

The ethical concept of fairness has recently been applied in machine learning (ML) settings to describe a wide range of constraints and objectives. When considering the relevance of ethical concepts to subset selection problems, the…

Artificial Intelligence · Computer Science 2020-02-11 Margaret Mitchell , Dylan Baker , Nyalleng Moorosi , Emily Denton , Ben Hutchinson , Alex Hanna , Timnit Gebru , Jamie Morgenstern

Lack of diversity in data collection has caused significant failures in machine learning (ML) applications. While ML developers perform post-collection interventions, these are time intensive and rarely comprehensive. Thus, new methods to…

Human-Computer Interaction · Computer Science 2023-08-01 Aspen Hopkins , Fred Hohman , Luca Zappella , Xavier Suau Cuadros , Dominik Moritz

Data is a crucial component of machine learning. The field is reliant on data to train, validate, and test models. With increased technical capabilities, machine learning research has boomed in both academic and industry settings, and one…

Computer Vision and Pattern Recognition · Computer Science 2021-09-20 Morgan Klaus Scheuerman , Emily Denton , Alex Hanna

Although diversity in NLP datasets has received growing attention, the question of how to measure it remains largely underexplored. This opinion paper examines the conceptual and methodological challenges of measuring data diversity and…

Computation and Language · Computer Science 2025-09-23 Dong Nguyen , Esther Ploeger

A recent study has shown that large-scale visual datasets are very biased: they can be easily classified by modern neural networks. However, the concrete forms of bias among these datasets remain unclear. In this study, we propose a…

Computer Vision and Pattern Recognition · Computer Science 2024-12-04 Boya Zeng , Yida Yin , Zhuang Liu

Imitation learning from large multi-task demonstration datasets has emerged as a promising path for building generally-capable robots. As a result, 1000s of hours have been spent on building such large-scale datasets around the globe.…

Datasets have played a foundational role in the advancement of machine learning research. They form the basis for the models we design and deploy, as well as our primary medium for benchmarking and evaluation. Furthermore, the ways in which…

Machine Learning · Computer Science 2021-11-16 Amandalynne Paullada , Inioluwa Deborah Raji , Emily M. Bender , Emily Denton , Alex Hanna

We introduce dataset multiplicity, a way to study how inaccuracies, uncertainty, and social bias in training datasets impact test-time predictions. The dataset multiplicity framework asks a counterfactual question of what the set of…

Machine Learning · Computer Science 2023-04-24 Anna P. Meyer , Aws Albarghouthi , Loris D'Antoni

This paper introduces the MERIT Dataset, a multimodal (text + image + layout) fully labeled dataset within the context of school reports. Comprising over 400 labels and 33k samples, the MERIT Dataset is a valuable resource for training…

Artificial Intelligence · Computer Science 2026-03-04 I. de Rodrigo , A. Sanchez-Cuadrado , J. Boal , A. J. Lopez-Lopez

Since its beginning visual recognition research has tried to capture the huge variability of the visual world in several image collections. The number of available datasets is still progressively growing together with the amount of samples…

Computer Vision and Pattern Recognition · Computer Science 2014-02-25 Tatiana Tommasi , Tinne Tuytelaars , Barbara Caputo

Data representativity is crucial when drawing inference from data through machine learning models. Scholars have increased focus on unraveling the bias and fairness in models, also in relation to inherent biases in the input data. However,…

Machine Learning · Statistics 2023-02-06 Line H. Clemmensen , Rune D. Kjærsgaard

Neural networks are powerful models that solve a variety of complex real-world problems. However, the stochastic nature of training and large number of parameters in a typical neural model makes them difficult to evaluate via inspection.…

Machine Learning · Computer Science 2021-04-22 John Clemens

Data heterogeneity plays a pivotal role in determining the performance of machine learning (ML) systems. Traditional algorithms, which are typically designed to optimize average performance, often overlook the intrinsic diversity within…

Machine Learning · Computer Science 2025-06-03 Jiashuo Liu , Peng Cui

We identify "values" as actions that classifiers take that speak to open questions of significant social concern. Investigating a classifier's values builds on studies of social bias that uncover how classifiers participate in social…

Computers and Society · Computer Science 2024-02-08 Will Penman , Joshua Babu , Abhinaya Raghunathan

In computer vision, a prevailing method for quantifying dataset bias is to train a model to distinguish between datasets. High classification accuracy is then interpreted as evidence of meaningful semantic differences. This approach assumes…

Computer Vision and Pattern Recognition · Computer Science 2026-04-16 Amir Hossein Saleknia , Mohammad Sabokrou
‹ Prev 1 2 3 10 Next ›