中文
相关论文

相关论文: Stochastic dynamic programming under recursive Eps…

200 篇论文

We propose a new stochastic primal-dual optimization algorithm for planning in a large discounted Markov decision process with a generative model and linear function approximation. Assuming that the feature map approximately satisfies…

机器学习 · 计算机科学 2023-02-01 Gergely Neu , Nneka Okolo

In this paper, we study the existence and uniqueness of solutions to a class of non-Lipschitz G-BSDEs and the corresponding stochastic recursive optimal control problem. More precisely, we suppose that the generator of G-BSDE is uniformly…

最优化与控制 · 数学 2026-04-14 Wei He , Qiangjun Tang

We study a discrete-time consumption-based capital asset pricing model under expectations-based reference-dependent preferences. More precisely, we consider an endowment economy populated by a representative agent who derives utility from…

数理金融 · 定量金融 2024-01-24 Luca De Gennaro Aquino , Xuedong He , Moris Simon Strub , Yuting Yang

In temporal difference (TD) learning, off-policy sampling is known to be more practical than on-policy sampling, and by decoupling learning from data collection, it enables data reuse. It is known that policy evaluation (including…

机器学习 · 计算机科学 2021-06-25 Zaiwei Chen , Siva Theja Maguluri , Sanjay Shakkottai , Karthikeyan Shanmugam

We present a simple and robust strategy for the selection of sampling points in Uncertainty Quantification. The goal is to achieve the fastest possible convergence in the cumulative distribution function of a stochastic output of interest.…

计算物理 · 物理学 2017-05-08 Enrico Camporeale , Ashutosh Agnihotri , Casper Rutjes

This paper is concerned with a class of stochastic optimization problems defined on a Banach space with almost sure conic-type constraints. For this class of problems, we investigate the consistency of optimal values and solutions…

最优化与控制 · 数学 2026-03-11 Caroline Geiersbach , Johannes Milz

The purpose of this note is to prove the existence of a randomized mechanism, a social decision scheme (SDS), with desirable fairness, efficiency, and strategyproofness properties unmatched by all known SDSs. In particular, we disprove a…

计算机科学与博弈论 · 计算机科学 2014-11-27 Florian Brandl

A major goal in Algorithmic Game Theory is to justify equilibrium concepts from an algorithmic and complexity perspective. One appealing approach is to identify natural distributed algorithms that converge quickly to an equilibrium. This…

计算机科学与博弈论 · 计算机科学 2018-06-14 Yun Kuen Cheung , Richard Cole , Yixin Tao

We describe a new approach for managing aleatoric uncertainty in the Reinforcement Learning (RL) paradigm. Instead of selecting actions according to a single statistic, we propose a distributional method based on the second-order stochastic…

机器学习 · 计算机科学 2020-10-08 John D. Martin , Michal Lyskawinski , Xiaohu Li , Brendan Englot

This work studies discrete-time discounted Markov decision processes with continuous state and action spaces and addresses the inverse problem of inferring a cost function from observed optimal behavior. We first consider the case in which…

最优化与控制 · 数学 2024-05-27 Angeliki Kamoutsi , Peter Schmitt-Förster , Tobias Sutter , Volkan Cevher , John Lygeros

We study the problem of aggregating distributions, such as budget proposals, into a collective distribution. An ideal aggregation mechanism would be Pareto efficient, strategyproof, and fair. Most previous work assumes that agents evaluate…

理论经济学 · 经济学 2025-12-24 Felix Brandt , Matthias Greger , Erel Segal-Halevi , Warut Suksompong

In this paper, we investigate the robust optimal reinsurance,investment,and internal surplus distribution (i.e., consumption) problem for an insurer with Epstein-Zin recursive preferences in an incomplete market. It is assumed that the…

最优化与控制 · 数学 2026-05-19 Junyi Guo , Jianxuan Li , Qianqian Zhou

We establish a new Bernstein-type deviation inequality for general (non-reversible) discrete-time Markov chains via an elementary approach. More robust than existing works in the literature, our result only requires the Markov chain to…

概率论 · 数学 2025-10-07 De Huang , Xiangyuan Li

We establish sharp large deviation principles for cumulative rewards associated with a discrete-time renewal model, supposing that each renewal involves a broad-sense reward taking values in a real separable Banach space. The framework we…

概率论 · 数学 2023-04-24 Marco Zamparo

We study the convergence of stochastic fixed point iterations in the consistent case (in the sense of Butnariu and Fl{\aa}m (1995)) in several different settings, under decreasingly restrictive regularity assumptions of the fixed point…

最优化与控制 · 数学 2020-03-26 Neal Hermer , D. Russell Luke , Anja Sturm

We present a distributional approach to theoretical analyses of reinforcement learning algorithms for constant step-sizes. We demonstrate its effectiveness by presenting simple and unified proofs of convergence for a variety of…

机器学习 · 计算机科学 2020-03-30 Philip Amortila , Doina Precup , Prakash Panangaden , Marc G. Bellemare

We study the convergence of random function iterations for finding an invariant measure of the corresponding Markov operator. We call the problem of finding such an invariant measure the stochastic fixed point problem. This generalizes…

最优化与控制 · 数学 2024-04-16 Neal Hermer , D. Russell Luke , Anja Sturm

We study a finite-inventory risk-sensitive market making problem in which a dealer controls bid and ask quotes, faces Brownian midprice risk, and receives liquidity-taking orders through point processes with quote-dependent intensities. The…

交易与市场微观结构 · 定量金融 2026-05-26 Tenghan Zhong

This paper is devoted to the study of acceleration methods for an inequality constrained convex optimization problem by using Lyapunov functions. We first approximate such a problem as an unconstrained optimization problem by employing the…

最优化与控制 · 数学 2024-11-25 Juan Liu , Nan-Jing Huang , Xian-Jun Long , Xue-song Li

Operator splitting techniques have recently gained popularity in convex optimization problems arising in various control fields. Being fixed-point iterations of nonexpansive operators, such methods suffer many well known downsides, which…

最优化与控制 · 数学 2020-04-01 Andreas Themelis , Panagiotis Patrinos