Related papers: Optimal run time for an EPQ model with scrap, rewo…
High update-to-data (UTD) ratio algorithms in reinforcement learning (RL) improve sample efficiency but incur high computational costs, limiting real-world scalability. We propose Offline Stabilization Phases for Efficient Q-Learning…
Motivated by applications where impatience is pervasive and evaluation times are uncertain, we study a selection model where options may expire at an unknown point in time and evaluation times are stochastic. Initially, the decision-maker…
Efficient energy management is essential for reliable and sustainable microgrid operation amid increasing renewable integration. In this paper, an imitation learning-based framework to approximate mixed-integer Economic Model Predictive…
In control applications there is often a compromise that needs to be made with regards to the complexity and performance of the controller and the computational resources that are available. For instance, the typical hardware platform in…
Scheduled batch jobs have been widely used on the asynchronous computing platforms to execute various enterprise applications, including the scheduled notifications and the candidate pre-computation for the modern recommender systems. It is…
The great majority of engineered products are subject to thermo-mechanical loads which vary with the product environment during the various phases of its life-cycle (machining, assembly, intended service use...). Those load variations may…
This work presents a suboptimality study of a particular model predictive control with a stage cost shaping based on the ideas of reinforcement learning. The focus of the suboptimality study is to derive quantities relating the…
Time-reversal symmetry breaking and entropy production are universal features of nonequilibrium phenomena. Despite its importance in the physics of active and living systems, the entropy production of systems with many degrees of freedom…
We present a mathematical model for optimizing breakaway strategies in competitive cycling, balancing power expenditure, aerodynamic drag, and crashing. Our framework incorporates probabilistic crash dynamics, allowing a cyclist's risk…
In the aftermath of the global financial crisis, much attention has been paid to investigating the appropriateness of the current practice of default risk modeling in banking, finance and insurance industries. A recent empirical study by…
We study expected runtimes for quantum programs. Inspired by recent work on probabilistic programs, we first define expected runtime as a generalisation of quantum weakest precondition. Then, we show that the expected runtime of a quantum…
Optimal execution is an important problem faced by any trader. Most solutions are based on the assumption of constant market impact, while liquidity is known to be dynamic. Moreover, models with time-varying liquidity typically assume that…
Seasonal climate variations affect electricity demand, which in turn affects month-to-month electricity planning and operations. Electricity system planning at the monthly timescale can be improved by adapting climate forecasts to estimate…
A high-speed train needs high-level maintenance when its accumulated running mileage or time reaches predefined threshold. The date of delivering an Electric Multiple Unit (EMU) train to maintenance ranges within a time window rather than…
A re-entrant manufacturing system producing a large number of items and involving many steps can be approximately modeled by a hyperbolic partial differential equation (PDE) according to mass conservation law with respect to a continuous…
The estimation of project completion time is to be repeated several times in the project planning phase to reach the optimal tradeoff between time, cost, and quality. Estimation procedures provide either an interval or a point estimate. The…
Given the ease of creating synthetic data from machine learning models, new models can be potentially trained on synthetic data generated by previous models. This recursive training process raises concerns about the long-term impact on…
A crucial problem in reinforcement learning is learning the optimal policy. We study this in tabular infinite-horizon discounted Markov decision processes under the online setting. The existing algorithms either fail to achieve regret…
Inventory models with imperfect quality items are studied by researchers in past two decades. Till now none of them have considered the effect of substitutions to cope up with shortage and avoid lost sales. This paper presents an EOQ…
When optimizing problems with uncertain parameter values in a linear objective, decision-focused learning enables end-to-end learning of these values. We are interested in a stochastic scheduling problem, in which processing times are…