中文
相关论文

相关论文: Planning in entropy-regularized Markov decision pr…

200 篇论文

This paper is focused on the study of entropic regularization in optimal transport as a smoothing method for Wasserstein estimators, through the prism of the classical tradeoff between approximation and estimation errors in statistics.…

机器学习 · 统计学 2024-10-30 Jérémie Bigot , Paul Freulon , Boris P. Hejblum , Arthur Leclaire

We explore the use of policy approximations to reduce the computational cost of learning Nash equilibria in zero-sum stochastic games. We propose a new Q-learning type algorithm that uses a sequence of entropy-regularized soft policies to…

机器学习 · 计算机科学 2021-06-29 Yue Guan , Qifan Zhang , Panagiotis Tsiotras

We consider the approximation of expectations with respect to the distribution of a latent Markov process given noisy measurements. This is known as the smoothing problem and is often approached with particle and Markov chain Monte Carlo…

统计计算 · 统计学 2019-02-06 Lawrence Middleton , George Deligiannidis , Arnaud Doucet , Pierre E. Jacob

This note summarizes the optimization formulations used in the study of Markov decision processes. We consider both the discounted and undiscounted processes under the standard and the entropy-regularized settings. For each setting, we…

最优化与控制 · 数学 2020-12-18 Lexing Ying , Yuhua Zhu

Nash equilibrium is a popular solution concept for solving imperfect-information games in practice. However, it has a major drawback: it does not preclude suboptimal play in branches of the game tree that are not reached in equilibrium.…

计算机科学与博弈论 · 计算机科学 2017-05-29 Christian Kroer , Gabriele Farina , Tuomas Sandholm

We want to introduce another smoothing approach by treating each geometric element as a player in a game: a quest for the best element quality. In other words, each player has the goal of becoming as regular as possible. The set of…

计算机科学与博弈论 · 计算机科学 2020-10-13 Dimitris Vartziotis , Doris Bohnet , Benjamin Himpel

In this paper, we focus on activating only a few sensors, among many available, to estimate the state of a stochastic process of interest. This problem is important in applications such as target tracking and simultaneous localization and…

系统与控制 · 计算机科学 2016-09-28 Vasileios Tzoumas , Nikolay A. Atanasov , Ali Jadbabaie , George J. Pappas

We propose and study a general framework for regularized Markov decision processes (MDPs) where the goal is to find an optimal policy that maximizes the expected discounted total reward plus a policy regularization term. The extant…

机器学习 · 统计学 2019-10-22 Xiang Li , Wenhao Yang , Zhihua Zhang

We study the problem of scheduling sensors in a resource-constrained linear dynamical system, where the objective is to select a small subset of sensors from a large network to perform the state estimation task. We formulate this problem as…

系统与控制 · 计算机科学 2018-04-05 Abolfazl Hashemi , Mahsa Ghasemi , Haris Vikalo , Ufuk Topcu

We consider a class of hierarchical noncooperative $N$-player games where the $i$th player solves a parametrized stochastic mathematical program with equilibrium constraints (MPEC) with the caveat that the implicit form of the $i$th…

最优化与控制 · 数学 2022-02-23 Shisheng Cui , Uday V. Shanbhag

In this work, we consider learning over multitask graphs, where each agent aims to estimate its own parameter vector. Although agents seek distinct objectives, collaboration among them can be beneficial in scenarios where relationships…

机器学习 · 计算机科学 2025-09-23 Yara Zgheib , Luca Calatroni , Marc Antonini , Roula Nassif

Large language models (LLMs) have shown promise in performing complex multi-step reasoning, yet they continue to struggle with mathematical reasoning, often making systematic errors. A promising solution is reinforcement learning (RL)…

机器学习 · 计算机科学 2025-09-22 Hanning Zhang , Pengcheng Wang , Shizhe Diao , Yong Lin , Rui Pan , Hanze Dong , Dylan Zhang , Pavlo Molchanov , Tong Zhang

We develop a general framework for state estimation in systems modeled with noise-polluted continuous time dynamics and discrete time noisy measurements. Our approach is based on maximum likelihood estimation and employs the calculus of…

最优化与控制 · 数学 2026-01-16 Griffin M. Kearney , Makan Fardad

We present the first finite-sample analysis of policy evaluation in robust average-reward Markov Decision Processes (MDPs). Prior work in this setting have established only asymptotic convergence guarantees, leaving open the question of…

机器学习 · 统计学 2025-12-11 Yang Xu , Washim Uddin Mondal , Vaneet Aggarwal

In this paper, we study the problem of estimating the state of a dynamic state-space system where the output is subject to quantization. We compare some classical approaches and a new development in the literature to obtain the filtering…

系统与控制 · 电气工程与系统科学 2021-12-16 Angel L. Cedeño , Ricardo Albornoz , Boris I. Godoy , Rodrigo Carvajal , Juan C. Agüero

Planning problems where effects of actions are non-deterministic can be modeled as Markov decision processes. Planning problems are usually goal-directed. This paper proposes several techniques for exploiting the goal-directedness to…

人工智能 · 计算机科学 2013-02-08 Nevin Lianwen Zhang , Weihong Zhang

This paper considers robust Markov decision processes under parametric transition distributions. We assume that the true transition distribution is uniquely specified by some parametric distribution, and explicitly enforce that the…

最优化与控制 · 数学 2022-11-24 Ben Black , Trivikram Dokka , Christopher Kirkbride

Ranking models primarily focus on modeling the relative order of predictions while often neglecting the significance of the accuracy of their absolute values. However, accurate absolute values are essential for certain downstream tasks,…

信息检索 · 计算机科学 2025-04-22 Yimeng Bai , Shunyu Zhang , Yang Zhang , Hu Liu , Wentian Bao , Enyun Yu , Fuli Feng , Wenwu Ou

In this paper is proposed a novel incremental iterative Gauss-Newton-Markov-Kalman filter method for state estimation of dynamic models given noisy measurements. The mathematical formulation of the proposed filter is based on the…

最优化与控制 · 数学 2019-09-17 Bojana Rosic

A central task of artificial intelligence is the design of artificial agents that act towards specified goals in partially observed environments. Since such environments frequently include interaction over time with other agents with their…

计算机科学与博弈论 · 计算机科学 2012-05-14 Miroslav Dudik , Geoffrey Gordon