English
Related papers

Related papers: Entropy-Regularized Certainty-Equivalent Bellman P…

200 papers

This paper characterizes differentiable and subgame Markov perfect equilibria in a continuous time intertemporal decision problem with non-constant discounting. Capturing the idea of non commitment by letting the commitment period being…

Optimization and Control · Mathematics 2008-08-29 Ivar Ekeland , Ali Lazrak

In this work, we propose a soft covering problem for fully quantum channels using relative entropy as a criterion for operator closeness. We establish covering lemmas by deriving one-shot bounds on the achievable rates in terms of smooth…

Information Theory · Computer Science 2026-02-24 Xingyi He , S. Sandeep Pradhan

We introduce a method for approximating viscosity solutions of stationary degenerate elliptic Hamilton--Jacobi--Bellman equations on bounded domains arising in stochastic exit-time control. Viscosity enforcement is formulated as a min--max…

Optimization and Control · Mathematics 2026-05-18 Alen E. Golpashin , Gokul Puthumanaillam , Melkior Ornik , Bruce A. Conway

In this paper, we present a novel algorithm named synchronous integral Q-learning, which is based on synchronous policy iteration, to solve the continuous-time infinite horizon optimal control problems of input-affine system dynamics. The…

Systems and Control · Electrical Eng. & Systems 2021-05-20 Lei Guo , Han Zhao

We consider a discrete-time dividend payout problem with risk sensitive shareholders. It is assumed that they are equipped with a risk aversion coefficient and construct their discounted payoff with the help of the exponential premium…

Probability · Mathematics 2017-03-08 Nicole Bäuerle , Anna Jaśkiewicz

Q-value iteration (Q-VI) is usually analyzed through the \(\gamma\)-contraction of the Bellman operator. This argument proves convergence to \(Q^*\), but it gives only a coarse account of when the induced greedy policy becomes optimal. We…

Optimization and Control · Mathematics 2026-05-06 Donghwan Lee

In this paper, we consider finite model approximations of a large class of static and dynamic team problems where these models are constructed through uniform quantization of the observation and action spaces of the agents. The strategies…

Optimization and Control · Mathematics 2016-01-05 Naci Saldi , Serdar Yüksel , Tamás Linder

Wasserstein policy gradient (WPG) is a policy optimization method for reinforcement learning (RL) that exploits the optimal-transport geometry of action distributions. For the entropy-regularized RL objective, WPG evolves each…

Machine Learning · Computer Science 2026-05-27 Zhaoyu Zhu , Rui Gao , Shuang Li

We describe an approximate dynamic programming approach to compute lower bounds on the optimal value function for a discrete time, continuous space, infinite horizon setting. The approach iteratively constructs a family of lower bounding…

Systems and Control · Electrical Eng. & Systems 2024-12-20 Paul N. Beuchat , Joseph Warrington , John Lygeros

Inventory models with lost sales and large lead times have traditionally been considered intractable due to the curse of dimensionality. Recently, Goldberg and co-authors laid the foundations for a new approach to solving these models, by…

Probability · Mathematics 2016-04-21 Linwei Xin , David A. Goldberg

A classical problem in ergodic continuous time control consists of studying the limit behavior of the optimal value of a discounted cost functional with infinite horizon as the discount factor $\lambda$ tends to zero. In the literature,…

Optimization and Control · Mathematics 2024-01-23 Piermarco Cannarsa , Stephane Gaubert , Cristian Mendico , Marc Quincampoix

We introduce the entropic measure transform (EMT) problem for a general process and prove the existence of a unique optimal measure characterizing the solution. The density process of the optimal measure is characterized using a…

Mathematical Finance · Quantitative Finance 2019-02-22 Renjie Wang , Cody Hyndman , Anastasis Kratsios

Establishing robust policies is essential to counter attacks or disturbances affecting deep reinforcement learning (DRL) agents. Recent studies explore state-adversarial robustness and suggest the potential lack of an optimal robust policy…

Machine Learning · Computer Science 2024-06-24 Haoran Li , Zicheng Zhang , Wang Luo , Congying Han , Yudong Hu , Tiande Guo , Shichen Liao

Existing value function approximation methods have been successfully used in many applications, but they often lack useful a priori error bounds. We propose a new approximate bilinear programming formulation of value function approximation,…

Artificial Intelligence · Computer Science 2010-06-15 Marek Petrik , Shlomo Zilberstein

The entropic risk measure is widely used in high-stakes decision-making across economics, management science, finance, and safety-critical control systems because it captures tail risks associated with uncertain losses. However, when data…

Optimization and Control · Mathematics 2026-01-05 Utsav Sadana , Erick Delage , Angelos Georghiou

We introduce Reliable Policy Iteration (RPI) and Conservative RPI (CRPI), variants of Policy Iteration (PI) and Conservative PI (CPI), that retain tabular guarantees under function approximation. RPI uses a novel Bellman-constrained…

Machine Learning · Computer Science 2026-04-03 S. R. Eshwar , Gugan Thoppe , Ananyabrata Barua , Aditya Gopalan , Gal Dalal

In this paper, a convex optimization-based method is proposed for numerically solving dynamic programs in continuous state and action spaces. The key idea is to approximate the output of the Bellman operator at a particular state by the…

Optimization and Control · Mathematics 2020-10-23 Insoon Yang

The dramatic increase of autonomous systems subject to variable environments has given rise to the pressing need to consider risk in both the synthesis and verification of policies for these systems. This paper aims to address a few…

Artificial Intelligence · Computer Science 2022-04-22 Prithvi Akella , Anushri Dixit , Mohamadreza Ahmadi , Joel W. Burdick , Aaron D. Ames

We consider a singular control problem with regime switching that arises in problems of optimal investment decisions of cash-constrained firms. The value function is proved to be the unique viscosity solution of the associated…

Computational Finance · Quantitative Finance 2016-10-07 Erwan Pierre , Stéphane Villeneuve , Xavier Warin

In this work we address the problem of finding feasible policies for Constrained Markov Decision Processes under probability one constraints. We argue that stationary policies are not sufficient for solving this problem, and that a rich…

Machine Learning · Computer Science 2023-02-14 Agustin Castellano , Hancheng Min , Juan Bazerque , Enrique Mallada
‹ Prev 1 8 9 10 Next ›