English

InvEvolve: Evolving White-Box Inventory Policies via Large Language Models with Performance Guarantees

Machine Learning 2026-05-12 v3 Artificial Intelligence

Abstract

We study how large language models can be used to evolve inventory policies in online, non-stationary environments. Our work is motivated by recent advances in LLM-based evolutionary search, such as AlphaEvolve, which demonstrates strong performance for static and highly structured problems such as mathematical discovery, but is not directly suited to online dynamic inventory settings. To this end, we propose InvEvolve, an end-to-end inventory policy evolution and inference framework grounded in confidence-interval-based certification. Built on a large language model trained via reinforcement learning, InvEvolve can process demand data together with additional numerical and textual features, and generates white-box inventory policies with statistical safety guarantees for deployment in future periods. We further introduce a unified theoretical model that connects training, inference, and deployment. This allows us to drive a lower bound on the probability that InvEvolve evolves a statistically safe and improved policy, and to characterize the multi-period performance gap relative to the oracle-safe benchmark. Tested on both synthetic data and real-world retail data, InvEvolve outperforms classical inventory policies and deep learning based methods. In canonical inventory settings, it evolves new policies that improve upon existing benchmarks.

Keywords

Cite

@article{arxiv.2605.00369,
  title  = {InvEvolve: Evolving White-Box Inventory Policies via Large Language Models with Performance Guarantees},
  author = {Chenyu Huang and Jianghao Lin and Zhengyang Tang and Bo Jiang and Ruoqing Jiang and Benyou Wang and Lai Wei},
  journal= {arXiv preprint arXiv:2605.00369},
  year   = {2026}
}
R2 v1 2026-07-01T12:44:44.731Z