English

SAGDA: Open-Source Synthetic Agriculture Data for Africa

Machine Learning 2025-06-17 v1 Machine Learning

Abstract

Data scarcity in African agriculture hampers machine learning (ML) model performance, limiting innovations in precision agriculture. The Synthetic Agriculture Data for Africa (SAGDA) library, a Python-based open-source toolkit, addresses this gap by generating, augmenting, and validating synthetic agricultural datasets. We present SAGDA's design and development practices, highlighting its core functions: generate, model, augment, validate, visualize, optimize, and simulate, as well as their roles in applications of ML for agriculture. Two use cases are detailed: yield prediction enhanced via data augmentation, and multi-objective NPK (nitrogen, phosphorus, potassium) fertilizer recommendation. We conclude with future plans for expanding SAGDA's capabilities, underscoring the vital role of open-source, data-driven practices for African agriculture.

Keywords

Cite

@article{arxiv.2506.13123,
  title  = {SAGDA: Open-Source Synthetic Agriculture Data for Africa},
  author = {Abdelghani Belgaid and Oumnia Ennaji},
  journal= {arXiv preprint arXiv:2506.13123},
  year   = {2025}
}