中文
相关论文

相关论文: Kernel-based diffusion approximated Markov decisio…

200 篇论文

We propose a principled kernel-based policy iteration algorithm to solve the continuous-state Markov Decision Processes (MDPs). In contrast to most decision-theoretic planning frameworks, which assume fully known state transition models, we…

机器人学 · 计算机科学 2020-06-04 Junhong Xu , Kai Yin , Lantao Liu

We consider policy evaluation in infinite-horizon discounted Markov decision problems (MDPs) with infinite spaces. We reformulate this task a compositional stochastic program with a function-valued decision variable that belongs to a…

最优化与控制 · 数学 2020-05-19 Alec Koppel , Garrett Warnell , Ethan Stump , Peter Stone , Alejandro Ribeiro

Motion planning under uncertainty for an autonomous system can be formulated as a Markov Decision Process with a continuous state space. In this paper, we propose a novel solution to this decision-theoretic planning problem that directly…

机器人学 · 计算机科学 2020-07-02 Junhong Xu , Kai Yin , Lantao Liu

With the advancement of autonomous driving, ensuring safety during motion planning and navigation is becoming more and more important. However, most end-to-end planning methods suffer from a lack of safety. This research addresses the…

人工智能 · 计算机科学 2024-07-18 Detian Chu , Linyuan Bai , Jianuo Huang , Zhenlong Fang , Peng Zhang , Wei Kang , Haifeng Lin

We introduce a distance between kernels based on the Wasserstein distances between their values, study its properties, and prove that it is a metric on an appropriately defined space of kernels. We also relate it to various modes of…

最优化与控制 · 数学 2024-01-29 Zhengqi Lin , Andrzej Ruszczynski

Diffusion models excel at sampling from complex, unnormalized distributions. In this work, we extend Maximum Entropy Reinforcement Learning (ME-RL) to diffusion processes, enabling sampling from the optimal policy trajectory distribution.…

机器学习 · 计算机科学 2026-05-28 Sebastian Sanokowski , Kaustubh Patil

We propose a new, nonparametric approach to learning and representing transition dynamics in Markov decision processes (MDPs), which can be combined easily with dynamic programming methods for policy optimisation and value estimation. This…

机器学习 · 计算机科学 2012-06-22 Steffen Grunewalder , Guy Lever , Luca Baldassarre , Massi Pontil , Arthur Gretton

This paper proposes a fully data-driven approach for optimal control of nonlinear control-affine systems represented by a stochastic diffusion. The focus is on the scenario where both the nonlinear dynamics and stage cost functions are…

最优化与控制 · 数学 2025-11-03 Nicolas Hoischen , Petar Bevanda , Stefan Sosnowski , Sandra Hirche , Boris Houska

We seek to learn an effective policy for a Markov Decision Process (MDP) with continuous states via Q-Learning. Given a set of basis functions over state action pairs we search for a corresponding set of linear weights that minimizes the…

机器学习 · 计算机科学 2013-09-27 Charles Tripp , Ross D. Shachter

Many real-world decision-making problems face the off-dynamics challenge: the agent learns a policy in a source domain and deploys it in a target domain with different state transitions. The distributionally robust Markov decision process…

机器学习 · 计算机科学 2025-05-26 Zhishuai Liu , Pan Xu

In this work, we consider an online robust Markov Decision Process (MDP) where we have the information of finitely many prototypes of the underlying transition kernel. We consider an adaptively updated ambiguity set of the prototypes and…

机器学习 · 计算机科学 2024-12-20 Shuo Sun , Meng Qi , Zuo-Jun Max Shen

A Robust Markov Decision Process (RMDP) is a sequential decision making model that accounts for uncertainty in the parameters of dynamic systems. This uncertainty introduces difficulties in learning an optimal policy, especially for…

人工智能 · 计算机科学 2017-03-08 Shirli Di-Castro Shashua , Shie Mannor

This work presents a multiscale framework to solve a class of stochastic optimal control problems in the context of robot motion planning and control in a complex environment. In order to handle complications resulting from a large decision…

机器人学 · 计算机科学 2017-03-14 Jung-Su Ha , Han-Lim Choi

Discrete time stochastic optimal control problems and Markov decision processes (MDPs) are fundamental models for sequential decision-making under uncertainty and as such provide the mathematical framework underlying reinforcement learning…

最优化与控制 · 数学 2025-07-01 Arnulf Jentzen , Konrad Kleinberg , Thomas Kruse

Markov Decision Processes (MDPs) have been used to formulate many decision-making problems in science and engineering. The objective is to synthesize the best decision (action selection) policies to maximize expected rewards (or minimize…

最优化与控制 · 数学 2015-07-07 Mahmoud El Chamie , Behcet Acikmese

Controllable diffusion generation often relies on various heuristics that are seemingly disconnected without a unified understanding. We bridge this gap with Diffusion Controller (DiffCon), a unified control-theoretic view that casts…

机器学习 · 计算机科学 2026-03-10 Tong Yang , Moonkyung Ryu , Chih-Wei Hsu , Guy Tennenholtz , Yuejie Chi , Craig Boutilier , Bo Dai

We study methods based on reproducing kernel Hilbert spaces for estimating the value function of an infinite-horizon discounted Markov reward process (MRP). We study a regularized form of the kernel least-squares temporal difference (LSTD)…

机器学习 · 统计学 2021-09-27 Yaqi Duan , Mengdi Wang , Martin J. Wainwright

A Markov decision process can be parameterized by a transition kernel and a reward function. Both play essential roles in the study of reinforcement learning as evidenced by their presence in the Bellman equations. In our inquiry of various…

机器学习 · 计算机科学 2023-09-04 Falcon Z. Dai

We develop algorithms with low regret for learning episodic Markov decision processes based on kernel approximation techniques. The algorithms are based on both the Upper Confidence Bound (UCB) as well as Posterior or Thompson Sampling…

机器学习 · 计算机科学 2019-11-06 Sayak Ray Chowdhury , Aditya Gopalan

Robust Markov decision processes (MDPs) allow to compute reliable solutions for dynamic decision problems whose evolution is modeled by rewards and partially-known transition probabilities. Unfortunately, accounting for uncertainty in the…

机器学习 · 计算机科学 2020-06-18 Chin Pang Ho , Marek Petrik , Wolfram Wiesemann
‹ 上一页 1 2 3 10 下一页 ›