English
Related papers

Related papers: Learning an Inventory Control Policy with General …

200 papers

Policy iteration is one of the classical frameworks of reinforcement learning, which requires a known initial stabilizing control. However, finding the initial stabilizing control depends on the known system model. To relax this requirement…

Systems and Control · Electrical Eng. & Systems 2025-03-20 Dongdong Li , Jiuxiang Dong

Problem definition: Supply chains are constantly evolving networks. Reinforcement learning is increasingly proposed as a solution to provide optimal control of these networks. Academic/practical: However, learning in continuously varying…

Systems and Control · Electrical Eng. & Systems 2023-12-27 Wan Wang , Haiyan Wang , Adam J. Sobey

We study a stylized dynamic assortment planning problem during a selling season of finite length $T$. At each time period, the seller offers an arriving customer an assortment of substitutable products and the customer makes the purchase…

Machine Learning · Statistics 2021-02-22 Xi Chen , Chao Shi , Yining Wang , Yuan Zhou

We consider the canonical periodic review lost sales inventory system with positive lead-times and stochastic i.i.d. demand under the average cost criterion. We introduce a new policy that places orders such that the expected inventory…

Probability · Mathematics 2024-01-17 Willem van Jaarsveld , Joachim Arts

E-commerce with major online retailers is changing the way people consume. The goal of increasing delivery speed while remaining cost-effective poses significant new challenges for supply chains as they race to satisfy the growing and…

Optimization and Control · Mathematics 2021-01-25 Adrien Rimélé , Philippe Grangier , Michel Gamache , Michel Gendreau , Louis-Martin Rousseau

We devise a control-theoretic reinforcement learning approach to support direct learning of the optimal policy. We establish various theoretical properties of our approach, such as convergence and optimality of our analog of the Bellman…

Machine Learning · Computer Science 2026-04-01 Weiqin Chen , Mark S. Squillante , Chai Wah Wu , Santiago Paternain

In this paper we are introducing a new reinforcement learning method for control problems in environments with delayed feedback. Specifically, our method employs stochastic planning, versus previous methods that used deterministic planning.…

Machine Learning · Computer Science 2024-02-02 Zhiyuan Yao , Ionut Florescu , Chihoon Lee

We propose a computationally efficient rollout-then-optimize method to improve a learned control policy at deployment time. A learned policy provides a nominal trajectory, which is refined online by a single Newton step implemented via a…

Optimization and Control · Mathematics 2026-04-13 Andrea Ghezzi , Rudolf Reiter , Katrin Baumgärtner , Alberto Bemporad , Moritz Diehl

In this paper, we present a novel model to characterize individual tendencies in repeated decision-making scenarios, with the goal of designing model-based control strategies that promote virtuous choices amidst social and external…

Systems and Control · Electrical Eng. & Systems 2025-03-06 Chiara Ravazzi , Valentina Breschi , Paolo Frasca , Fabrizio Dabbene , Mara Tanelli

A multi-class single-server queueing model with finite buffers, in which scheduling and admission of customers are subject to control, is studied in the moderate deviation heavy traffic regime. A risk-sensitive cost set over a finite time…

Probability · Mathematics 2018-05-02 Rami Atar , Asaf Cohen

We consider a network inventory system motivated by one-way, on-demand vehicle sharing services. Under uncertain and correlated network demand, the service operator periodically repositions vehicles to match a fixed supply with spatial…

Machine Learning · Statistics 2025-10-20 Hansheng Jiang , Chunlin Sun , Zuo-Jun Max Shen

Inventory-policy comparisons are often difficult to interpret because performance depends on the evaluation contract as much as on the policy itself. Differences in topology, demand regime, information access, feasibility constraints,…

Machine Learning · Computer Science 2026-05-13 Reza Barati , Qinmin Vivian Hu

When learning policies for real-world domains, two important questions arise: (i) how to efficiently use pre-collected off-policy, non-optimal behavior data; and (ii) how to mediate among different competing objectives and constraints. We…

Machine Learning · Computer Science 2019-03-22 Hoang M. Le , Cameron Voloshin , Yisong Yue

We examine the dynamics of the bid and ask queues of a limit order book and their relationship with the intensity of trade arrivals. In particular, we study the probability of price movements and trade arrivals as a function of the quote…

Trading and Market Microstructure · Quantitative Finance 2013-12-03 Alexander Lipton , Umberto Pesavento , Michael G Sotiropoulos

A new method for controlling harmonic generation, in the framework of quantum optimal control theory (QOCT), is developed. The problem is formulated in the frequency domain using a new maximization functional. The relaxation method is used…

Quantum Physics · Physics 2012-08-28 Ido Schaefer

Offline Reinforcement learning is commonly used for sequential decision-making in domains such as healthcare and education, where the rewards are known and the transition dynamics $T$ must be estimated on the basis of batch data. A key…

Machine Learning · Computer Science 2023-08-10 Leo Benac , Sonali Parbhoo , Finale Doshi-Velez

A common pipeline in learning-based control is to iteratively estimate a model of system dynamics, and apply a trajectory optimization algorithm - e.g.~$\mathtt{iLQR}$ - on the learned model to minimize a target cost. This paper conducts a…

Machine Learning · Computer Science 2023-05-17 Daniel Pfrommer , Max Simchowitz , Tyler Westenbroek , Nikolai Matni , Stephen Tu

Modelling of contact-rich tasks is challenging and cannot be entirely solved using classical control approaches due to the difficulty of constructing an analytic description of the contact dynamics. Additionally, in a manipulation task like…

Robotics · Computer Science 2019-09-27 Ioanna Mitsioni , Yiannis Karayiannidis , Johannes A. Stork , Danica Kragic

We study the problem of policy repair for learning-based control policies in safety-critical settings. We consider an architecture where a high-performance learning-based control policy (e.g. one trained as a neural network) is paired with…

Artificial Intelligence · Computer Science 2020-08-19 Weichao Zhou , Ruihan Gao , BaekGyu Kim , Eunsuk Kang , Wenchao Li

In retail warehouses, returned products are typically placed in an intermediate storage until a decision regarding further shipment to stores is made. The longer products are held in storage, the higher the inefficiency and costs of the…

Machine Learning · Computer Science 2025-01-27 Pascal Linden , Nathalie Paul , Tim Wirtz , Stefan Wrobel