To cite this paper use one of the standards below:
Lot-sizing decisions under volatile demand, capacity limits, setup interactions, and service requirements are difficult to address with static models alone. Reinforcement learning (RL) offers an adaptive alternative, but evidence is dispersed across lot-sizing, inventory control, production planning, and scheduling. Following PRISMA guidelines, this article presents a structured review using Scopus, citation chaining, and comparative bibliometric and content analyses of an initial portfolio of 317 publications and a refined core of 31 articles. Results indicate a transition from tabular and agent-based learning toward deep RL, especially policy-gradient, actor-critic, proximal policy optimization, and multi-agent methods. The refined literature is predominantly simulation-based and increasingly addresses complex stochastic settings, while standardized benchmarks, comparisons with exact methods, explainability, feasibility preservation, reproducibility, and industrial validation remain limited. These gaps define the main requirements for developing dependable RL-based lot-sizing decision-support tools.
With nearly 200,000 papers published, Galoá empowers scholars to share and discover cutting-edge research through our streamlined and accessible academic publishing platform.
Learn more about our products:
This proceedings is identified by a DOI , for use in citations or bibliographic references. Attention: this is not a DOI for the paper and as such cannot be used in Lattes to identify a particular work.
Check the link "How to cite" in the paper's page, to see how to properly cite the paper