English

Data-driven inventory management for new products: An adjusted Dyna-$Q$ approach with transfer learning

Machine Learning 2025-06-10 v4 Artificial Intelligence Computational Engineering, Finance, and Science

Abstract

In this paper, we propose a novel reinforcement learning algorithm for inventory management of newly launched products with no historical demand information. The algorithm follows the classic Dyna-QQ structure, balancing the model-free and model-based approaches, while accelerating the training process of Dyna-QQ and mitigating the model discrepancy generated by the model-based feedback. Based on the idea of transfer learning, warm-start information from the demand data of existing similar products can be incorporated into the algorithm to further stabilize the early-stage training and reduce the variance of the estimated optimal policy. Our approach is validated through a case study of bakery inventory management with real data. The adjusted Dyna-QQ shows up to a 23.7\% reduction in average daily cost compared with QQ-learning, and up to a 77.5\% reduction in training time within the same horizon compared with classic Dyna-QQ. By using transfer learning, it can be found that the adjusted Dyna-QQ has the lowest total cost, lowest variance in total cost, and relatively low shortage percentages among all the benchmarking algorithms under a 30-day testing.

Cite

@article{arxiv.2501.08109,
  title  = {Data-driven inventory management for new products: An adjusted Dyna-$Q$ approach with transfer learning},
  author = {Xinye Qu and Longxiao Liu and Wenjie Huang},
  journal= {arXiv preprint arXiv:2501.08109},
  year   = {2025}
}

Comments

7 pages, 3 figures

R2 v1 2026-06-28T21:05:54.311Z