中文
相关论文

相关论文: Dynamic allocation indices for restless projects a…

200 篇论文

We consider reinforcement learning (RL) in episodic MDPs with adversarial full-information reward feedback and unknown fixed transition kernels. We propose two model-free policy optimization algorithms, POWER and POWER++, and establish…

机器学习 · 计算机科学 2020-07-02 Yingjie Fei , Zhuoran Yang , Zhaoran Wang , Qiaomin Xie

Efficient exact algorithms for Discrete Optimization (DO) rely heavily on strong primal and dual bounds. Relaxed Decision Diagrams (DDs) provide a versatile mechanism for deriving such dual bounds by compactly over-approximating the…

人工智能 · 计算机科学 2025-12-18 Mohsen Nafar , Michael Römer , Lin Xie

Reinforcement learning (RL)-based quadrotor control policies have achieved impressive performance in tasks such as fast navigation in cluttered environments and drone racing, where the focus is on speed and agility. However, in several…

机器人学 · 计算机科学 2026-05-20 Fausto Mauricio Lagos Suarez , Akshit Saradagi , Vidya Sumathy , George Nikolakopoulos

Robust Markov decision processes (MDPs) allow to compute reliable solutions for dynamic decision problems whose evolution is modeled by rewards and partially-known transition probabilities. Unfortunately, accounting for uncertainty in the…

机器学习 · 计算机科学 2020-06-18 Chin Pang Ho , Marek Petrik , Wolfram Wiesemann

Rather than augmenting rewards with penalties for undesired behavior, Constrained Partially Observable Markov Decision Processes (CPOMDPs) plan safely by imposing inviolable hard constraint value budgets. Previous work performing online…

人工智能 · 计算机科学 2022-12-26 Arec Jamgochian , Anthony Corso , Mykel J. Kochenderfer

Many existing studies on mixed-criticality (MC) scheduling assume that low-criticality budgets for high-criticality applications are known apriori. These budgets are primarily used as guidance to determine when the scheduler should switch…

分布式、并行与集群计算 · 计算机科学 2020-03-19 Xiaozhe Gu , Arvind Easwaran

We consider a system with a local cache connected to a backend server and an end user population. A set of contents are stored at the the server where they continuously get updated. The local cache keeps copies, potentially stale, of a…

网络与互联网体系结构 · 计算机科学 2025-04-10 Ankita Koley , Chandramani Singh

We are interested in risk constraints for infinite horizon discrete time Markov decision processes (MDPs). Starting with average reward MDPs, we show that increasing concave stochastic dominance constraints on the empirical distribution of…

最优化与控制 · 数学 2012-06-21 William B. Haskell , Rahul Jain

Most reinforcement-learning (RL) controllers used in continuous control are architecturally centralized: observations are compressed into a single latent state from which both value estimates and actions are produced. Biological control…

机器学习 · 计算机科学 2026-04-27 Anne E. Staples

We study the optimal design of stealthy attacks against partially observed linear control systems. We first propose a novel likelihood-based detection mechanism derived from the innovation process, based on which we quantify stealthiness…

最优化与控制 · 数学 2026-05-12 Haosheng Zhou , Ruimeng Hu

We consider the problem of service placement at the network edge, in which a decision maker has to choose between $N$ services to host at the edge to satisfy the demands of customers. Our goal is to design adaptive algorithms to minimize…

网络与互联网体系结构 · 计算机科学 2021-01-15 Guojun Xiong , Rahul Singh , Jian Li

The dynamic allocation problem, also known as the `multi-armed bandit' problem, simulates a situation in which an agent is faced with a tradeoff between actions that yield an immediate reward and actions whose benefits can only be perceived…

概率论 · 数学 2026-02-03 Christopher Wang

We study a problem of information gathering in a social network with dynamically available sources and time varying quality of information. We formulate this problem as a restless multi-armed bandit (RMAB). In this problem, information…

系统与控制 · 计算机科学 2018-01-22 Varun Mehta , Rahul Meshram , Kesav Kaza , S. N. Merchant

Autonomous agents are limited in their ability to observe the world state. Partially observable Markov decision processes (POMDPs) formally model the problem of planning under world state uncertainty, but POMDPs with continuous actions and…

机器人学 · 计算机科学 2020-07-08 Dicong Qiu , Yibiao Zhao , Chris L. Baker

This paper presents an algorithm to apply nonlinear control design approaches in the case of stochastic systems with partial state observation. Deterministic nonlinear control approaches are formulated under the assumption of full state…

系统与控制 · 电气工程与系统科学 2023-09-19 Mohammad S. Ramadan , Mohammad Alsuwaidan , Ahmed Atallah , Sylvia Herbert

We consider a stochastic, dynamic job scheduling problem, formulated as a queueing control problem, in which a single server processes jobs of different types that arrive according to independent Poisson processes. The problem is defined on…

最优化与控制 · 数学 2025-09-09 Dongnuan Tian , Rob Shone

We introduce a class of distributed control policies for networks of discrete-time linear systems with polytopic additive disturbances. The objective is to restrict the network-level state and controls to user-specified polyhedral sets for…

系统与控制 · 计算机科学 2017-09-29 Sadra Sadraddini , Calin Belta

We explore online learning in episodic loop-free Markov decision processes on non-stationary environments (changing losses and probability transitions). Our focus is on the Concave Utility Reinforcement Learning problem (CURL), an extension…

机器学习 · 计算机科学 2024-05-31 Bianca Marin Moreno , Margaux Brégère , Pierre Gaillard , Nadia Oudjane

We study episodic reinforcement learning (RL) in non-stationary linear kernel Markov decision processes (MDPs). In this setting, both the reward function and the transition kernel are linear with respect to the given feature maps and are…

机器学习 · 计算机科学 2024-12-24 Han Zhong , Zhongren Chen , Zhuoran Yang , Zhaoran Wang , Csaba Szepesvári

We propose and analyze a model-based bootstrap for transition kernels in finite controlled Markov chains (CMCs) with possibly nonstationary or history-dependent control policies, a setting that arises naturally in offline reinforcement…

机器学习 · 统计学 2026-05-13 Ziwei Su , Imon Banerjee , Diego Klabjan
‹ 上一页 1 8 9 10 下一页 ›