English
Related papers

Related papers: Value iteration for approximate dynamic programmin…

200 papers

This paper studies stochastic optimization problems and associated Bellman equations in formats that allow for reduced dimensionality of the cost-to-go functions. In particular, we study stochastic control problems in the…

Optimization and Control · Mathematics 2025-05-20 Teemu Pennanen , Ari-Pekka Perkkiö

In a Markovian framework, we consider the problem of finding the minimal initial value of a controlled process allowing to reach a stochastic target with a given level of expected loss. This question arises typically in approximate hedging…

Optimization and Control · Mathematics 2017-04-06 Géraldine Bouveret , Jean-François Chassagneux

Stochastic time-varying optimization is an integral part of learning in which the shape of the function changes over time in a non-deterministic manner. This paper considers multiple models of stochastic time variation and analyzes the…

Optimization and Control · Mathematics 2023-02-23 Ali Yekkehkhany , Han Feng , Donghao Ying , Javad Lavaei

This paper considers an infinite-horizon Markov decision process (MDP) that allows for general non-exponential discount functions, in both discrete and continuous time. Due to the inherent time inconsistency, we look for a randomized…

Optimization and Control · Mathematics 2024-12-10 Erhan Bayraktar , Yu-Jui Huang , Zhenhua Wang , Zhou Zhou

We consider linear programming (LP) problems in infinite dimensional spaces that are in general computationally intractable. Under suitable assumptions, we develop an approximation bridge from the infinite-dimensional LP to tractable finite…

Optimization and Control · Mathematics 2017-02-22 Peyman Mohajerin Esfahani , Tobias Sutter , Daniel Kuhn , John Lygeros

Primal-dual splitting involving proximity operators in order to be able to find some approximation to the minimizer for a general form of Tikhonov type functional is in the focus of this work. This approximation is produced by a pair of…

Numerical Analysis · Mathematics 2019-03-19 Erdem Altuntac

In this paper we present an algorithm for pricing barrier options in one-dimensional Markov models. The approach rests on the construction of an approximating continuous-time Markov chain that closely follows the dynamics of the given…

Pricing of Securities · Quantitative Finance 2015-03-13 Aleksandar Mijatovic , Martijn Pistorius

We consider a dynamic programming problem with arbitrary state space and bounded rewards. Is it possible to define in an unique way a limit value for the problem, where the "patience" of the decision-maker tends to infinity ? We consider,…

Optimization and Control · Mathematics 2013-01-04 Jérôme Renault

In this paper, we present a new set-valued Lagrange multiplier theorem for constrained convex set-valued optimization problems. We introduce the novel concept of Lagrange process. This concept is a natural extension of the classical concept…

Optimization and Control · Mathematics 2024-01-19 Fernando García-Castaño , M. A. Melguizo Padial

In this work, we deal with an iteration method for approximating a fixed point of a contraction mapping using the Mann's algorithm under functional random errors. We first show its almost complete convergence to the fixed point by mean of…

Probability · Mathematics 2017-01-24 Bahia Barache , Idir Arab , Abdelnasser Dahmani

In this paper, it is shown that Bermudan option pricing based on either the r\'eduite (in a one-dimensional setting: piecewise harmonic interpolation) or cubature -- is sensible from an economic vantage point: Any sequence of thus-computed…

Probability · Mathematics 2007-05-23 Frederik S. Herzberg

Partially observable Markov decision processes (POMDPs) provide an elegant mathematical framework for modeling complex decision and planning problems in stochastic domains in which states of the system are observable only indirectly, via a…

Artificial Intelligence · Computer Science 2011-06-02 M. Hauskrecht

We formally verify executable algorithms for solving Markov decision processes (MDPs) in the interactive theorem prover Isabelle/HOL. We build on existing formalizations of probability theory to analyze the expected total reward criterion…

Artificial Intelligence · Computer Science 2023-03-09 Maximilian Schäfeller , Mohammad Abdulaziz

It is well known that iterates of quasi-compact operators converge towards a spectral projection, whereas the explicit construction of the limiting operator is in general hard to obtain. Here, we show a simple method to explicitly construct…

Functional Analysis · Mathematics 2017-06-05 Johannes Nagler

We present a convex-concave reformulation of the reversible Markov chain estimation problem and outline an efficient numerical scheme for the solution of the resulting problem based on a primal-dual interior point method for monotone…

Data Analysis, Statistics and Probability · Physics 2016-03-08 Benjamin Trendelkamp-Schroer , Hao Wu , Frank Noe

In many branches of engineering, Banach contraction mapping theorem is employed to establish the convergence of certain deterministic algorithms. Randomized versions of these algorithms have been developed that have proved useful in…

Probability · Mathematics 2023-09-25 Abhishek Gupta , Rahul Jain , Peter Glynn

This paper shows the usefulness of the Perov contraction theorem, which is a generalization of the classical Banach contraction theorem, for solving Markov dynamic programming problems. When the reward function is unbounded, combining an…

Optimization and Control · Mathematics 2024-05-06 Alexis Akira Toda

Designing efficient learning algorithms with complexity guarantees for Markov decision processes (MDPs) with large or continuous state and action spaces remains a fundamental challenge. We address this challenge for entropy-regularized MDPs…

Machine Learning · Computer Science 2025-06-05 Matthieu Meunier , Christoph Reisinger , Yufei Zhang

The problem of determining the (least) fixpoint of (higher-dimensional) functions over the non-negative reals frequently occurs when dealing with systems endowed with a quantitative semantics. We focus on the situation in which the…

Logic in Computer Science · Computer Science 2026-01-23 Paolo Baldan , Sebastian Gurke , Barbara König , Florian Wittbold

One of the most widely used methods for solving average cost MDP problems is the value iteration method. This method, however, is often computationally impractical and restricted in size of solvable MDP problems. We propose acceleration…

Optimization and Control · Mathematics 2008-06-03 Oleksandr Shlakhter , Chi-Guhn Lee