中文
相关论文

相关论文: Inventory Management with Partially Observed Nonst…

200 篇论文

We consider the problem of maximizing expected utility for a power investor who can allocate his wealth in a stock, a defaultable security, and a money market account. The dynamics of these security prices are governed by geometric Brownian…

投资组合管理 · 定量金融 2014-06-04 Agostino Capponi , Jose Enrique Figueroa Lopez , Andrea Pascucci

We study a production-inventory system with two customer classes with different priorities which are admitted to the system following a flexible admission control scheme. The inventory management is according to a base stock policy and…

概率论 · 数学 2023-03-21 Sonja Otten , Hans Daduna

To identify a stationary action profile for a population of competitive agents, each executing private strategies, we introduce a novel active-learning scheme where a centralized external observer (or entity) can probe the agents' reactions…

系统与控制 · 电气工程与系统科学 2024-10-10 Filippo Fabiani , Alberto Bemporad

We cast episodic Markov decision process (MDP) planning as Bayesian inference over policies. A policy is treated as the latent variable and is assigned an unnormalized probability of optimality that is monotone in its expected return,…

机器学习 · 计算机科学 2026-04-14 David Tolpin

In this paper we complete and extend our previous work on stochastic control applied to high frequency market-making with inventory constraints and directional bets. Our new model admits several state variables (e.g. market spread,…

交易与市场微观结构 · 定量金融 2013-04-03 Pietro Fodra , Mauricio Labadie

In this paper we propose a framework towards achieving two intertwined objectives: (i) equipping reinforcement learning with active exploration and deliberate information gathering, such that it regulates state and parameter uncertainties…

机器学习 · 计算机科学 2024-09-10 Mohammad S. Ramadan , Mahmoud A. Hayajnh , Michael T. Tolley , Kyriakos G. Vamvoudakis

In this paper we study the stochastic control problem of partially observed (multi-dimensional) stochastic system driven by both Brownian motions and fractional Brownian motions. In the absence of the powerful tool of Girsanov…

最优化与控制 · 数学 2023-08-22 Yueyang Zheng , Yaozhong Hu

This paper implements the Deep Deterministic Policy Gradient (DDPG) algorithm for computing optimal policies for partially observable single-product periodic review inventory control problems with setup costs and backorders. The decision…

最优化与控制 · 数学 2025-07-29 Eugene Feinberg , Jefferson Huang , Pavlo Kasyanov , Thomas O'Neill

We consider infinite-horizon $\gamma$-discounted Markov Decision Processes, for which it is known that there exists a stationary optimal policy. We consider the algorithm Value Iteration and the sequence of policies $\pi_1,...,\pi_k$ it…

人工智能 · 计算机科学 2012-04-02 Bruno Scherrer

In applications of offline reinforcement learning to observational data, such as in healthcare or education, a general concern is that observed actions might be affected by unobserved factors, inducing confounding and biasing estimates…

机器学习 · 计算机科学 2023-03-24 Andrew Bennett , Nathan Kallus

We consider an inventory system whose state is modeled by a L\'{e}vy process. There are two types of costs--the running costs and the inventory control costs. The running costs (also known as the holding/penalty costs) are incurred…

最优化与控制 · 数学 2016-09-02 Jinbiao Wu , Haolin Feng , Dacheng Yao

A tenet of reinforcement learning is that the agent always observes rewards. However, this is not true in many realistic settings, e.g., a human observer may not always be available to provide rewards, sensors may be limited or…

机器学习 · 计算机科学 2026-03-24 Alireza Kazemipour , Simone Parisi , Matthew E. Taylor , Michael Bowling

Continuous control and planning remains a major challenge in robotics and machine learning. Neuroscience offers the possibility of learning from animal brains that implement highly successful controllers, but it is unclear how to relate an…

人工智能 · 计算机科学 2019-08-14 Saurabh Daptardar , Paul Schrater , Xaq Pitkow

We consider the problem of diagnosis where a set of simple observations are used to infer a potentially complex hidden hypothesis. Finding the optimal subset of observations is intractable in general, thus we focus on the problem of active…

人工智能 · 计算机科学 2017-07-12 Yewen Pu , Leslie P Kaelbling , Armando Solar-Lezama

Inference for partially observed Markov process models has been a longstanding methodological challenge with many scientific and engineering applications. Iterated filtering algorithms maximize the likelihood function for partially observed…

统计理论 · 数学 2012-11-26 Edward L. Ionides , Anindya Bhadra , Yves Atchadé , Aaron King

We present an online stochastic model predictive control framework for demand charge management for a grid-connected consumer with attached electrical energy storage. The consumer we consider must satisfy an inflexible but stochastic…

系统与控制 · 电气工程与系统科学 2020-07-07 Benjamin Flamm , Guillermo Ramos , Annika Eichler , John Lygeros

The problem of pure exploration in Markov decision processes has been cast as maximizing the entropy over the state distribution induced by the agent's policy, an objective that has been extensively studied. However, little attention has…

机器学习 · 计算机科学 2024-06-19 Riccardo Zamboni , Duilio Cirino , Marcello Restelli , Mirco Mutti

Computational level explanations based on optimal feedback control with signal-dependent noise have been able to account for a vast array of phenomena in human sensorimotor behavior. However, commonly a cost function needs to be assumed for…

机器学习 · 计算机科学 2021-10-22 Matthias Schultheis , Dominik Straub , Constantin A. Rothkopf

We study the problem of synthesizing a controller that maximizes the entropy of a partially observable Markov decision process (POMDP) subject to a constraint on the expected total reward. Such a controller minimizes the predictability of a…

最优化与控制 · 数学 2019-09-16 Michael Hibbard , Yagiz Savas , Bo Wu , Takashi Tanaka , Ufuk Topcu

In most real-world reinforcement learning applications, state information is only partially observable, which breaks the Markov decision process assumption and leads to inferior performance for algorithms that conflate observations with…

机器学习 · 计算机科学 2024-06-12 Hongming Zhang , Tongzheng Ren , Chenjun Xiao , Dale Schuurmans , Bo Dai