English
Related papers

Related papers: Time-inconsistent mean-field stopping problems: A …

200 papers

We consider the game-theoretic approach to time-inconsistent stopping of a one-dimensional diffusion where the time-inconsistency is due to the presence of a non-exponential (weighted) discount function. In particular, we study (weak)…

Probability · Mathematics 2022-07-01 Andi Bodnariu , Sören Christensen , Kristoffer Lindensjö

In this paper, a sparse Markov decision process (MDP) with novel causal sparse Tsallis entropy regularization is proposed.The proposed policy regularization induces a sparse and multi-modal optimal policy distribution of a sparse MDP. The…

Machine Learning · Computer Science 2017-10-16 Kyungjae Lee , Sungjoon Choi , Songhwai Oh

This note describes sufficient conditions under which total-cost and average-cost Markov decision processes (MDPs) with general state and action spaces, and with weakly continuous transition probabilities, can be reduced to discounted MDPs.…

Optimization and Control · Mathematics 2017-11-21 Eugene A. Feinberg , Jefferson Huang

We study the synthesis of a policy in a Markov decision process (MDP) following which an agent reaches a target state in the MDP while minimizing its total discounted cost. The problem combines a reachability criterion with a discounted…

Optimization and Control · Mathematics 2021-03-18 Yagiz Savas , Christos K. Verginis , Michael Hibbard , Ufuk Topcu

The problem of constrained Markov decision process is considered. An agent aims to maximize the expected accumulated discounted reward subject to multiple constraints on its costs (the number of constraints is relatively small). A new dual…

Optimization and Control · Mathematics 2022-10-21 Egor Gladin , Maksim Lavrik-Karmazin , Karina Zainullina , Varvara Rudenko , Alexander Gasnikov , Martin Takáč

We consider a single-server queue where interarrival and service times depend linearly and randomly on customer waiting times, and establish a sample-path moderate deviation principle (MDP) for the waiting time process. The waiting times…

Probability · Mathematics 2025-11-03 Chang Feng , John J. Hasenbein , Guodong Pang

In this paper, we consider the discounted continuous-time Markov decision process (CTMDP) with a lower bounding function. In this model, the negative part of each cost rate is bounded by the drift function, say $w$, whereas the positive…

Optimization and Control · Mathematics 2016-12-05 Xin Guo , Alexey Piunovskiy , Yi Zhang

Reinforcement learning in non-stationary environments is challenging due to abrupt and unpredictable changes in dynamics, often causing traditional algorithms to fail to converge. However, in many real-world cases, non-stationarity has some…

Machine Learning · Computer Science 2025-03-25 Mohsen Amiri , Sindri Magnússon

We study a specific class of finite-horizon mean field optimal stopping problems by means of the dynamic programming approach. In particular, we consider problems where the state process is not affected by the stopping time. Such problems…

Optimization and Control · Mathematics 2025-03-07 Andrea Cosso , Laura Perelli

Algorithms developed under stationary Markov Decision Processes (MDPs) often face challenges in non-stationary environments, and infinite-horizon formulations may not directly apply to finite-horizon tasks. To address these limitations, we…

Machine Learning · Computer Science 2025-12-03 Zhizuo Chen , Theodore T. Allen

We consider a mean-field control problem in which admissible controls are required to be adapted to the common noise filtration. The main objective is to show how the mean-field control problem can be approximates by time consistent…

Optimization and Control · Mathematics 2025-09-19 Bruno Bouchard , Xiaolu Tan

Multi-agent reinforcement learning methods have shown remarkable potential in solving complex multi-agent problems but mostly lack theoretical guarantees. Recently, mean field control and mean field games have been established as a…

Machine Learning · Computer Science 2021-12-20 Kai Cui , Anam Tahir , Mark Sinzger , Heinz Koeppl

The analysis of Temporal Difference (TD) learning in the average-reward setting faces notable theoretical difficulties because the Bellman operator is not contractive with respect to any norm. This complicates standard analyses of…

Machine Learning · Computer Science 2026-05-05 Haoxing Tian , Zaiwei Chen , Ioannis Ch. Paschalidis , Alex Olshevsky

Motivated by applications in risk-sensitive reinforcement learning, we study mean-variance optimization in a discounted reward Markov Decision Process (MDP). Specifically, we analyze a Temporal Difference (TD) learning algorithm with linear…

Machine Learning · Computer Science 2025-03-13 Tejaram Sangadi , L. A. Prashanth , Krishna Jagannathan

Robust Markov decision processes (MDPs) aim to handle changing or partially known system dynamics. To solve them, one typically resorts to robust optimization methods. However, this significantly increases computational complexity and…

Machine Learning · Computer Science 2023-03-14 Esther Derman , Yevgeniy Men , Matthieu Geist , Shie Mannor

This note re-visits the rolling-horizon control approach to the problem of a Markov decision process (MDP) with infinite-horizon discounted expected reward criterion. Distinguished from the classical value-iteration approach, we develop an…

Optimization and Control · Mathematics 2022-06-07 Hyeong Soo Chang

In this paper, which is a continuation of the previously published discrete time paper we develop a theory for continuous time stochastic control problems which, in various ways, are time inconsistent in the sense that they do not admit a…

Optimization and Control · Mathematics 2016-12-13 Tomas Björk , Mariana Khapko , Agatha Murgoci

This paper focuses on a class of continuous-time controlled Markov chains with time-inconsistent and distribution-dependent cost functional (in some appropriate sense). A new definition of time-inconsistent distribution-dependent…

Optimization and Control · Mathematics 2019-09-26 Hongwei Mei , George Yin

We study global optimization of non-convex functions through optimal control theory. Our main result establishes that (quasi-)optimal trajectories of a discounted control problem converge globally and practically asymptotically to the set…

Optimization and Control · Mathematics 2025-11-17 Yuyang Huang , Dante Kalise , Hicham Kouhkouh

We provide a characterization of an optimal stopping time for a class of finite horizon time-inconsistent optimal stopping problems (OSPs) of mean-field type, adapted to the Brownian filtration, including those related to mean-field…

Probability · Mathematics 2023-07-20 Boualem Djehiche , Mattia Martini