English

Feature Encodings for Gradient Boosting with Automunge

Machine Learning 2022-10-27 v2

Abstract

Automunge is a tabular preprocessing library that encodes dataframes for supervised learning. When selecting a default feature encoding strategy for gradient boosted learning, one may consider metrics of training duration and achieved predictive performance associated with the feature representations. Automunge offers a default of binarization for categoric features and z-score normalization for numeric. The presented study sought to validate those defaults by way of benchmarking on a series of diverse data sets by encoding variations with tuned gradient boosted learning. We found that on average our chosen defaults were top performers both from a tuning duration and a model performance standpoint. Another key finding was that one hot encoding did not perform in a manner consistent with suitability to serve as a categoric default in comparison to categoric binarization. We present here these and further benchmarks.

Cite

@article{arxiv.2209.12309,
  title  = {Feature Encodings for Gradient Boosting with Automunge},
  author = {Nicholas J. Teague},
  journal= {arXiv preprint arXiv:2209.12309},
  year   = {2022}
}

Comments

10 pages, 4 figures, preprint

R2 v1 2026-06-28T02:03:32.846Z