English
Related papers

Related papers: Entropy-Regularized Certainty-Equivalent Bellman P…

200 papers

The maximum entropy principle is a powerful tool for solving underdetermined inverse problems. This paper considers the problem of discretizing a continuous distribution, which arises in various applied fields. We obtain the approximating…

Numerical Analysis · Mathematics 2020-08-05 Ken'ichiro Tanaka , Alexis Akira Toda

Recent advances in continuous-time optimal stopping have been driven by entropy-regularized formulations of randomized stopping problems, with most existing approaches relying on partial differential equation methods. In this paper, we…

Computational Finance · Quantitative Finance 2026-02-23 Daniel Chee , Noufel Frikha , Libo Li

According to the entropy accumulation theorem, proving the unconditional security of a device-independent quantum key distribution protocol reduces to deriving tradeoff functions, i.e., bounds on the single-round von Neumann entropy of the…

Quantum Physics · Physics 2022-10-28 Michele Masini , Stefano Pironio , Erik Woodhead

We propose empirical dynamic programming algorithms for Markov decision processes (MDPs). In these algorithms, the exact expectation in the Bellman operator in classical value iteration is replaced by an empirical estimate to get `empirical…

Optimization and Control · Mathematics 2013-11-26 William B. Haskell , Rahul Jain , Dileep Kalathil

We consider the minimization over probability measures of the expected value of a random variable, regularized by relative entropy with respect to a given probability distribution. In the general setting we provide a complete…

Probability · Mathematics 2020-09-29 Joris Bierkens , Hilbert J. Kappen

Offline reinforcement learning promises policy improvement from logged interaction data alone, yet state-of-the-art algorithms remain vulnerable to value over-estimation and to violations of domain knowledge such as monotonicity or…

Systems and Control · Electrical Eng. & Systems 2025-06-18 Ali Baheri

We introduce a framework for approximate dynamic programming that we apply to discrete time chains on $\mathbb{Z}_+^d$ with countable action sets. Our approach is grounded in the approximation of the (controlled) chain's generator by that…

Optimization and Control · Mathematics 2018-04-16 Anton Braverman , Itai Gurvich , Junfei Huang

Maximum entropy deep reinforcement learning (RL) methods have been demonstrated on a range of challenging continuous tasks. However, existing methods either suffer from severe instability when training on large off-policy data or cannot…

Machine Learning · Computer Science 2019-09-10 Wenjie Shi , Shiji Song , Cheng Wu

Maximum entropy reinforcement learning (RL) methods have been successfully applied to a range of challenging sequential decision-making and control tasks. However, most of existing techniques are designed for discrete-time systems. As a…

Optimization and Control · Mathematics 2020-09-29 Jeongho Kim , Insoon Yang

Analyzing and controlling system entropy is a powerful tool for regulating predictability of control systems. Applications benefiting from such approaches range from reinforcement learning and data security to human-robot collaboration. In…

Systems and Control · Electrical Eng. & Systems 2026-03-06 Menno van Zutphen , Giannis Delimpaltadakis , Duarte J. Antunes

Efficient exploration is a central problem in reinforcement learning and is often formalized as maximizing the entropy of the state-action occupancy measure. While unconstrained maximum-entropy exploration is relatively well understood,…

Machine Learning · Computer Science 2026-05-01 Florian Wolf , Ilyas Fatkhullin , Niao He

Existing work on risk-sensitive reinforcement learning - both for symmetric and downside risk measures - has typically used direct Monte-Carlo estimation of policy gradients. While this approach yields unbiased gradient estimates, it also…

Machine Learning · Computer Science 2020-07-09 Thomas Spooner , Rahul Savani

Merton portfolio management problem is studied in this paper within a stochastic volatility, non constant time discount rate, and power utility framework. This problem is time inconsistent and the way out of this predicament is to consider…

Portfolio Management · Quantitative Finance 2024-02-09 Oumar Mbodji , Traian A. Pirvu

Entropy augmented to reward is known to soften the greedy argmax policy to softmax policy. Entropy augmentation is reformulated and leads to a motivation to introduce an additional entropy term to the objective function in the form of…

Machine Learning · Computer Science 2020-06-08 Donghoon Lee

The analysis of Temporal Difference (TD) learning in the average-reward setting faces notable theoretical difficulties because the Bellman operator is not contractive with respect to any norm. This complicates standard analyses of…

Machine Learning · Computer Science 2026-05-05 Haoxing Tian , Zaiwei Chen , Ioannis Ch. Paschalidis , Alex Olshevsky

The empirical risk minimization (ERM) problem with relative entropy regularization (ERM-RER) is investigated under the assumption that the reference measure is a $\sigma$-finite measure, and not necessarily a probability measure. Under this…

Statistics Theory · Mathematics 2024-04-09 Samir M. Perlaza , Gaetan Bisson , Iñaki Esnaola , Alain Jean-Marie , Stefano Rini

We study an OTC FX market-making problem, built on the Avellaneda-Stoikov tradition, in which a dealer streams size-dependent quotes on a discrete ladder and manages inventory risk over a finite horizon under Poisson arrivals of trade…

Risk Management · Quantitative Finance 2026-03-10 Alexander Barzykin

This paper concerns rollout and certainty-equivalent rollout policies for stochastic shortest path problems with absorbing terminal states. The main result provides a direct non-asymptotic performance certificate for a fixed rollout policy:…

Optimization and Control · Mathematics 2026-05-25 Anders Hansson , Bo Wahlberg

We investigate the portfolio execution problem under a framework in which volatility and liquidity are both uncertain. In our model, we assume that a multidimensional Markovian stochastic factor drives both of them. Moreover, we model…

Mathematical Finance · Quantitative Finance 2023-08-08 Max O. Souza , Yuri Thamsten

The entropy regularization is inspired by information entropy from machine learning and the ideas of exploration and exploitation in reinforcement learning, which appears in the control problem to design an approximating algorithm for the…

Optimization and Control · Mathematics 2024-11-21 Ziyue Chen , Qi Zhang