中文
相关论文

相关论文: Minimax-Regret Climate Policy with Deep Uncertaint…

200 篇论文

The parameters for a Markov Decision Process (MDP) often cannot be specified exactly. Uncertain MDPs (UMDPs) capture this model ambiguity by defining sets which the parameters belong to. Minimax regret has been proposed as an objective for…

人工智能 · 计算机科学 2023-02-14 Marc Rigter , Bruno Lacerda , Nick Hawes

We consider the setting of iterative learning control, or model-based policy learning in the presence of uncertain, time-varying dynamics. In this setting, we propose a new performance metric, planning regret, which replaces the standard…

机器学习 · 计算机科学 2021-03-01 Naman Agarwal , Elad Hazan , Anirudha Majumdar , Karan Singh

Optimizing a building's energy supply design is a task with multiple competing criteria, where not only monetary but also, for example, an environmental objective shall be taken into account. Moreover, when deciding which storages and…

最优化与控制 · 数学 2024-07-26 Elisabeth Halser , Elisabeth Finhold , Neele Leithäuser , Tobias Seidel , Karl-Heinz Küfer

It is common to use minimax rules to make decisions for planning when there is great uncertainty on what will happen in the future. Minimax regret is one popular version of this. We give an analysis of the behaviour of minimax rules in the…

最优化与控制 · 数学 2022-03-04 Edward Anderson , Stan Zachary

Large uncertainties in the energy transition urge decision-makers to develop low-regret strategies, i.e., strategies that perform well regardless of how the future unfolds. To address this challenge, we introduce a decision-support…

系统与控制 · 电气工程与系统科学 2025-05-20 Gabriel Wiest , Niklas Nolzen , Florian Baader , André Bardow , Stefano Moret

I study the problem of a decision maker choosing a policy which allocates treatment to a heterogeneous population on the basis of experimental data that includes only a subset of possible treatment values. The effects of new treatments are…

计量经济学 · 经济学 2025-07-17 Samuel Higbee

A crucial problem in reinforcement learning is learning the optimal policy. We study this in tabular infinite-horizon discounted Markov decision processes under the online setting. The existing algorithms either fail to achieve regret…

机器学习 · 计算机科学 2023-12-13 Xiang Ji , Gen Li

An important problem in sequential decision-making under uncertainty is to use limited data to compute a safe policy, i.e., a policy that is guaranteed to perform at least as well as a given baseline strategy. In this paper, we develop and…

机器学习 · 统计学 2016-07-14 Marek Petrik , Yinlam Chow , Mohammad Ghavamzadeh

While the importance of personalized policymaking is widely recognized, fully personalized implementation remains rare in practice, often due to legal, fairness or cost concerns. We study the problem of policy targeting for a regret-averse…

计量经济学 · 经济学 2026-04-07 Toru Kitagawa , Sokbae Lee , Chen Qiu

The exploration/exploitation trade-off is an inherent challenge in data-driven adaptive control. Though this trade-off has been studied for multi-armed bandits (MAB's) and reinforcement learning for linear systems; it is less well-studied…

最优化与控制 · 数学 2023-01-30 Ilgin Dogan , Zuo-Jun Max Shen , Anil Aswani

Least worst regret (and sometimes minimax) analysis are often used for decision making whenever it is difficult, or inappropriate, to attach probabilities to possible future scenarios. We show that, for each of these two approaches and…

最优化与控制 · 数学 2016-08-03 Stan Zachary

Threshold policies are decision rules that assign treatments based on whether an observable characteristic exceeds a certain threshold. They are widespread across multiple domains, including welfare programs, taxation, and clinical…

计量经济学 · 经济学 2025-04-08 Federico Crippa

This article improves the existing proven rates of regret decay in optimal policy estimation. We give a margin-free result showing that the regret decay for estimating a within-class optimal policy is second-order for empirical risk…

统计理论 · 数学 2017-04-24 Alexander Luedtke , Antoine Chambaz

In practical applications, data is used to make decisions in two steps: estimation and optimization. First, a machine learning model estimates parameters for a structural model relating decisions to outcomes. Second, a decision is chosen to…

最优化与控制 · 数学 2022-10-28 Samuel Tan , Peter I. Frazier

For decision making under uncertainty, min-max regret has been established as a popular methodology to find robust solutions. In this approach, we compare the performance of our solution against the best possible performance had we known…

最优化与控制 · 数学 2021-11-25 Marc Goerigk , Michael Hartisch

Markov decision processes (MDPs) are widely used in modeling decision making problems in stochastic environments. However, precise specification of the reward functions in MDPs is often very difficult. Recent approaches have focused on…

人工智能 · 计算机科学 2012-02-20 Eunsoo Oh , Kee-Eung Kim

We study agents acting in an unknown environment where the agent's goal is to find a robust policy. We consider robust policies as policies that achieve high cumulative rewards for all possible environments. To this end, we consider agents…

We consider estimation and control in linear time-varying dynamical systems from the perspective of regret minimization. Unlike most prior work in this area, we focus on the problem of designing causal estimators and controllers which…

机器学习 · 计算机科学 2021-06-24 Gautam Goel , Babak Hassibi

The specification of aMarkov decision process (MDP) can be difficult. Reward function specification is especially problematic; in practice, it is often cognitively complex and time-consuming for users to precisely specify rewards. This work…

人工智能 · 计算机科学 2012-05-14 Kevin Regan , Craig Boutilier

In unsupervised environment design, reinforcement learning agents are trained on environment configurations (levels) generated by an adversary that maximises some objective. Regret is a commonly used objective that theoretically results in…

‹ 上一页 1 2 3 10 下一页 ›