中文
相关论文

相关论文: Blind Dynamic Resource Allocation in Closed Networ…

200 篇论文

Policy optimization methods like Group Relative Policy Optimization (GRPO) and its variants have achieved strong results on mathematical reasoning and code generation tasks. Despite extensive exploration of reward processing strategies and…

机器学习 · 计算机科学 2026-02-05 Rui Yuan , Mykola Khandoga , Vinay Kumar Sankarapu

Entropy regularization has been widely used in policy optimization algorithms to enhance exploration and the robustness of the optimal control; however it also introduces an additional regularization bias. This work quantifies the impact of…

最优化与控制 · 数学 2025-03-25 Deven Sethi , David Šiška , Yufei Zhang

We consider a matching system with random arrivals of items of different types. The items wait in queues -- one per each item type -- until they are "matched." Each matching requires certain quantities of items of different types; after a…

概率论 · 数学 2018-04-24 Mohammadreza Nazari , Alexander L. Stolyar

This paper considers a cross-layer adaptive modulation system that is modeled as a Markov decision process (MDP). We study how to utilize the monotonicity of the optimal transmission policy to relieve the computational complexity of dynamic…

机器学习 · 统计学 2015-08-25 Ni Ding , Parastoo Sadeghi , Rodney A. Kennedy

Optimizing the assortment of products to display to customers is a key to increasing revenue for both offline and online retailers. To trade-off between exploring customers' preference and exploiting customers' choices learned from data, in…

机器学习 · 计算机科学 2022-04-25 Hongbin Zhang , Yu Yang , Feng Wu , Qixin Zhang

Recent advances in 3D fabrication have allowed handling the memory bottlenecks for modern data-intensive applications by bringing the computation closer to the memory, enabling Near Memory Processing (NMP). Memory Centric Networks (MCN) are…

硬件体系结构 · 计算机科学 2023-12-14 Shubhang Pandey , T G Venkatesh

Reinforcement learning (RL) is currently one of the most prominent methods for optimizing dynamical systems, with breakthrough results across various fields. The framework is based on the concept of a Markov decision process (MDP), leading…

最优化与控制 · 数学 2025-11-17 Rene Carmona , Mathieu Lauriere

We consider the problem faced by a service platform that needs to match limited supply with demand but also to learn the attributes of new users in order to match them better in the future. We introduce a benchmark model with heterogeneous…

机器学习 · 计算机科学 2020-08-07 Ramesh Johari , Vijay Kamble , Yash Kanoria

Model Predictive Control (MPC) is a well-established approach to solve infinite horizon optimal control problems. Since optimization over an infinite time horizon is generally infeasible, MPC determines a suboptimal feedback control by…

最优化与控制 · 数学 2022-10-26 Saskia Dietze , Martin A. Grepl

In this paper, we formulate an optimal ordering policy as a stochastic control problem where each firm decides the amount of input goods to order from their upstream suppliers based on the current inventory level of its output good. For…

最优化与控制 · 数学 2022-09-13 Jose I. Caiza , Ian Walter , Jitesh H. Panchal , Junjie Qin , Philip E. Pare

This paper considers a Markov decision model for profit maximization of a cloud computing service provider catering to customers submitting jobs with firm real-time random deadlines. Customers are charged on a per-job basis, receiving a…

最优化与控制 · 数学 2021-04-27 José Niño-Mora

We consider assortment and inventory planning problems with dynamic stockout-based substitution effects, and without replenishment, in two different settings: (1) Customers can see all available products when they arrive, a typical scenario…

最优化与控制 · 数学 2025-01-09 Shuo Sun , Rajan Udwani , Zuo-Jun Max Shen

This paper analyzes a service system modeled as a single-server queue, in which the service provider aims to dynamically maximize the expected revenue per unit of time. This is achieved by constructing a stochastic gradient descent…

最优化与控制 · 数学 2026-03-05 Shreehari Anand Bodas , Harsha Honnappa , Michel Mandjes , Liron Ravner

This paper targets at the problem of radio resource management for expected long-term delay-power tradeoff in vehicular communications. At each decision epoch, the road side unit observes the global network state, allocates channels and…

信号处理 · 电气工程与系统科学 2019-06-04 Xianfu Chen , Celimuge Wu , Honggang Zhang , Yan Zhang , Mehdi Bennis , Heli Vuojala

Memory-Bounded Dynamic Programming (MBDP) has proved extremely effective in solving decentralized POMDPs with large horizons. We generalize the algorithm and improve its scalability by reducing the complexity with respect to the number of…

人工智能 · 计算机科学 2012-06-26 Sven Seuken , Shlomo Zilberstein

Motivated by applications in online marketplaces such as ride-hailing, we study how strategic servers impact the system performance. We consider a discrete-time process in which, heterogeneous types of customers and servers arrive. Each…

最优化与控制 · 数学 2021-06-25 Sushil Mahavir Varma , Francisco Castro , Siva Theja Maguluri

We propose a new policy gradient method, named homotopic policy mirror descent (HPMD), for solving discounted, infinite horizon MDPs with finite state and action spaces. HPMD performs a mirror descent type policy update with an additional…

机器学习 · 计算机科学 2022-11-30 Yan Li , Guanghui Lan , Tuo Zhao

Pricing decisions are often made when market information is still poor. In turn, existing theoretical models often reason about the response of optimal prices to changing market characteristics without exploiting all available information…

最优化与控制 · 数学 2021-07-19 Stefanos Leonardos , Costis Melolidakis , Constandina Koki

This paper presents an efficient suboptimal model predictive control (MPC) algorithm for nonlinear switched systems subject to minimum dwell time constraints (MTC). While MTC are required for most physical systems due to stability, power…

最优化与控制 · 数学 2022-02-16 Yutao Chen , Mircea Lazar

The cost of the power distribution infrastructures is driven by the peak power encountered in the system. Therefore, the distribution network operators consider billing consumers behind a common transformer in the function of their peak…

系统与控制 · 电气工程与系统科学 2022-04-01 Wenqi Cai , Hossein N. Esfahani , Arash B. Kordabad , Sébastien Gros
‹ 上一页 1 8 9 10 下一页 ›