中文
相关论文

相关论文: Planning in entropy-regularized Markov decision pr…

200 篇论文

We study multi-agent general-sum Markov games with nonlinear function approximation. We focus on low-rank Markov games whose transition matrix admits a hidden low-rank structure on top of an unknown non-linear representation. The goal is to…

机器学习 · 计算机科学 2022-11-01 Chengzhuo Ni , Yuda Song , Xuezhou Zhang , Chi Jin , Mengdi Wang

Meta-heuristics are powerful tools for solving optimization problems whose structural properties are unknown or cannot be exploited algorithmically. We propose such a meta-heuristic for a large class of optimization problems over discrete…

离散数学 · 计算机科学 2021-06-22 Moritz Mühlenthaler , Alexander Raß , Manuel Schmitt , Rolf Wanka

We consider the task of estimating a structural model of dynamic decisions by a human agent based upon the observable history of implemented actions and visited states. This problem has an inherent nested structure: in the inner problem, an…

机器学习 · 计算机科学 2024-03-04 Siliang Zeng , Mingyi Hong , Alfredo Garcia

We propose a new stochastic primal-dual optimization algorithm for planning in a large discounted Markov decision process with a generative model and linear function approximation. Assuming that the feature map approximately satisfies…

机器学习 · 计算机科学 2023-02-01 Gergely Neu , Nneka Okolo

Constructing confidence intervals for the value of an (unknown) optimal treatment policy is a fundamental problem in causal inference. Insight into the optimal policy value can guide the development of reward-maximizing, individualized…

计量经济学 · 经济学 2026-04-01 Justin Whitehouse , Qizhao Chen , Morgane Austern , Vasilis Syrgkanis

We study infinite-horizon robust Markov decision processes (MDPs) on continuous state spaces with structured rectangular ambiguity set. The proposed ambiguity set falls within the convex hull of unknown generating kernels. We utilize the…

最优化与控制 · 数学 2026-05-28 Mengmeng Li , Yifan Hu , Daniel Kuhn , Yan Li

Entropy regularization is commonly used to improve policy optimization in reinforcement learning. It is believed to help with \emph{exploration} by encouraging the selection of more stochastic policies. In this work, we analyze this claim…

机器学习 · 计算机科学 2019-06-11 Zafarali Ahmed , Nicolas Le Roux , Mohammad Norouzi , Dale Schuurmans

We study entropy-regularized constrained Markov decision processes (CMDPs) under the soft-max parameterization, in which an agent aims to maximize the entropy-regularized value function while satisfying constraints on the expected total…

机器学习 · 计算机科学 2023-04-10 Donghao Ying , Yuhao Ding , Javad Lavaei

Model-based planners for partially observable problems must accommodate both model uncertainty during planning and goal uncertainty during objective inference. However, model-based planners may be brittle under these types of uncertainty…

Multi-time-scale stochastic approximation is an iterative algorithm for finding the fixed point of a set of $N$ coupled operators given their noisy samples. It has been observed that due to the coupling between the decision variables and…

最优化与控制 · 数学 2024-09-13 Sihan Zeng , Thinh T. Doan

We propose a novel polyhedral uncertainty set for robust optimization, termed the smooth uncertainty set, which captures dependencies of uncertain parameters by constraining their pairwise differences. The bounds on these differences may be…

最优化与控制 · 数学 2025-10-13 Noam Goldberg , Michael Poss , Shimrit Shtern

We propose to smooth out the calibration score, which measures how good a forecaster is, by combining nearby forecasts. While regular calibration can be guaranteed only by randomized forecasting procedures, we show that smooth calibration…

理论经济学 · 经济学 2022-10-14 Dean P. Foster , Sergiu Hart

Nonlinear model predictive control (NMPC) is a popular strategy for solving motion planning problems, including obstacle avoidance constraints, in autonomous driving applications. Non-smooth obstacle shapes, such as rectangles, introduce…

系统与控制 · 电气工程与系统科学 2024-03-05 Rudolf Reiter , Katrin Baumgärtner , Rien Quirynen , Moritz Diehl

Existing training criteria in automatic speech recognition(ASR) permit the model to freely explore more than one time alignments between the feature and label sequences. In this paper, we use entropy to measure a model's uncertainty, i.e.…

计算与语言 · 计算机科学 2022-12-26 Ehsan Variani , Ke Wu , David Rybach , Cyril Allauzen , Michael Riley

We present an efficient robust value iteration for \texttt{s}-rectangular robust Markov Decision Processes (MDPs) with a time complexity comparable to standard (non-robust) MDPs which is significantly faster than any existing method. We do…

机器学习 · 计算机科学 2023-02-01 Navdeep Kumar , Kfir Levy , Kaixin Wang , Shie Mannor

Quantum magic, or nonstabilizerness, provides a crucial characterization of quantum systems, regarding the classical simulability with stabilizer states. In this work, we propose a novel and efficient algorithm for computing stabilizer…

量子物理 · 物理学 2025-02-25 Zejun Liu , Bryan K. Clark

We investigate the problem of designing optimal classifiers in the strategic classification setting, where the classification is part of a game in which players can modify their features to attain a favorable classification outcome (while…

机器学习 · 计算机科学 2020-05-19 Mark Braverman , Sumegha Garg

In this paper, we settle the sampling complexity of solving discounted two-player turn-based zero-sum stochastic games up to polylogarithmic factors. Given a stochastic game with discount factor $\gamma\in(0,1)$ we provide an algorithm that…

机器学习 · 计算机科学 2019-08-30 Aaron Sidford , Mengdi Wang , Lin F. Yang , Yinyu Ye

This paper presents a new safety specification method that is robust against errors in the probability distribution of disturbances. Our proposed distributionally robust safe policy maximizes the probability of a system remaining in a…

最优化与控制 · 数学 2018-10-05 Insoon Yang

Probabilistic Circuits (PCs) are a promising avenue for probabilistic modeling. They combine advantages of probabilistic graphical models (PGMs) with those of neural networks (NNs). Crucially, however, they are tractable probabilistic…

机器学习 · 计算机科学 2021-06-07 Anji Liu , Guy Van den Broeck