English

The Lean Data Scientist: Recent Advances towards Overcoming the Data Bottleneck

Machine Learning 2022-11-16 v1 Artificial Intelligence

Abstract

Machine learning (ML) is revolutionizing the world, affecting almost every field of science and industry. Recent algorithms (in particular, deep networks) are increasingly data-hungry, requiring large datasets for training. Thus, the dominant paradigm in ML today involves constructing large, task-specific datasets. However, obtaining quality datasets of such magnitude proves to be a difficult challenge. A variety of methods have been proposed to address this data bottleneck problem, but they are scattered across different areas, and it is hard for a practitioner to keep up with the latest developments. In this work, we propose a taxonomy of these methods. Our goal is twofold: (1) We wish to raise the community's awareness of the methods that already exist and encourage more efficient use of resources, and (2) we hope that such a taxonomy will contribute to our understanding of the problem, inspiring novel ideas and strategies to replace current annotation-heavy approaches.

Keywords

Cite

@article{arxiv.2211.07959,
  title  = {The Lean Data Scientist: Recent Advances towards Overcoming the Data Bottleneck},
  author = {Chen Shani and Jonathan Zarecki and Dafna Shahaf},
  journal= {arXiv preprint arXiv:2211.07959},
  year   = {2022}
}
R2 v1 2026-06-28T05:55:41.484Z