Back to Research papers
Research paper index

InvEvolve: Evolving White-Box Inventory Policies via Large Language Models with Performance Guarantees

Chenyu Huang, Jianghao Lin, Zhengyang Tang, Bo Jiang, Ruoqing Jiang, Benyou Wang, Lai Wei

arXiv:2605.00369Published May 1, 2026Updated May 9, 20260 citations
  • cs.LG
  • cs.AI
  • reinforcement learning
  • policy

Abstract

We study how large language models can be used to evolve inventory policies in online, non-stationary environments. Our work is motivated by recent advances in LLM-based evolutionary search, such as AlphaEvolve, which demonstrates strong performance for static and highly structured problems such as mathematical discovery, but is not directly suited to online dynamic inventory settings. To this end, we propose InvEvolve, an end-to-end inventory policy evolution and inference framework grounded in confidence-interval-based certification. Built on a large language model trained via reinforcement learning, InvEvolve can process demand data together with additional numerical and textual features, and generates white-box inventory policies with statistical safety guarantees for deployment in future periods. We further introduce a unified theoretical model that connects training, inference, and deployment. This allows us to drive a lower bound on the probability that InvEvolve evolves a statistically safe and improved policy, and to characterize the multi-period performance gap relative to the oracle-safe benchmark. Tested on both synthetic data and real-world retail data, InvEvolve outperforms classical inventory policies and deep learning based methods. In canonical inventory settings, it evolves new policies that improve upon existing benchmarks.

Read the original paper

This page indexes public paper metadata. The manuscript remains with its original publisher and authors.