中文
相关论文

相关论文: Analysis and Evaluation of Baseline Manipulation i…

200 篇论文

In this paper, we show how a simulated Markov decision process (MDP) built by the so-called \emph{baseline} policies, can be used to compute a different policy, namely the \emph{simulated optimal} policy, for which the performance of this…

最优化与控制 · 数学 2014-10-13 Yinlam Chow , Mohammad Ghavamzadeh

Many applications -- including power systems, robotics, and economics -- involve a dynamical system interacting with a stochastic and hard-to-model environment. We adopt a reinforcement learning approach to control such systems.…

最优化与控制 · 数学 2025-08-26 Abed AlRahman Al Makdah , Oliver Kosut , Lalitha Sankar , Shaofeng Zou

It can be observed that the purchasing decision of an individual consumer in an electronic marketplace is determined by a set of factors, such as personal characteristics of the consumer, product pricing, minimum price-quantity combination…

计算机科学与博弈论 · 计算机科学 2024-06-27 Dipankar Das

Dynamic programming (DP) is a fundamental tool used across many engineering fields. The main goal of DP is to solve Bellman's optimality equations for a given Markov decision process (MDP). Standard methods like policy iteration exploit the…

人工智能 · 计算机科学 2025-07-30 Sergio Rozada , Samuel Rey , Gonzalo Mateos , Antonio G. Marques

Domain randomization (DR) is widely used in policy learning to improve robustness to modeling error, but remains underexplored in contact-rich sampling-based predictive control (SPC), where rollout quality is highly sensitive to…

机器人学 · 计算机科学 2026-05-06 Sergio A. Esteban , Junheng Li , Vince Kurtz , Aaron D. Ames

We develop a model for pricing, lead-time quotation and delay compensation in a Markovian make-to-order production or service system with strategic customers who exhibit risk aversion. Based on a concave utility function of their net…

最优化与控制 · 数学 2019-11-07 Myron Benioudakis , Apostolos Burnetas , George Ioannou

Technical debt is a metaphor that describes the long term effects of shortcuts taken in software development activities to achieve near term goals. In this study, we explore a new context of technical debt that relates to database…

软件工程 · 计算机科学 2018-01-26 Mashel Albarak , Muna Alrazgan , Rami Bahsoon

In this paper, a novel optimal control-based baseline function is presented for the policy gradient method in deep reinforcement learning (RL). The baseline is obtained by computing the value function of an optimal control problem, which is…

机器学习 · 计算机科学 2024-11-07 Xubo Lyu , Site Li , Seth Siriya , Ye Pu , Mo Chen

Uplift modeling aims to directly model the incremental impact of a treatment on an individual response. In this work, we address the problem from a new angle and reformulate it as a Markov Decision Process (MDP). We conducted extensive…

机器学习 · 计算机科学 2019-02-06 Chenchen Li , Xiang Yan , Xiaotie Deng , Yuan Qi , Wei Chu , Le Song , Junlong Qiao , Jianshan He , Junwu Xiong

Demand response is widely employed by today's data centers to reduce energy consumption in response to the increasing of electricity cost. To incentivize users of data centers participate in the demand response programs, i.e., breaking the…

分布式、并行与集群计算 · 计算机科学 2016-04-08 Yong Zhan , Du Xu , Hongfang Yu , Shui Yu

In this paper we develop an algorithm for peak load reduction to reduce the impact of increased air conditioner usage in a residential smart grid community. We develop Demand Response Management (DRM) plans that clearly spell out the…

系统与控制 · 计算机科学 2014-08-07 Yawar Ismail Khalid , Naveed Ul Hassan , Chau Yuen , Shisheng Huang

Price elasticity model (PEM) is an appealing and modest model for assessing the potential of flexible demand in DR. It measures the customers demand sensitivity through elasticity in relation to price variation. However, application of PEM…

系统与控制 · 电气工程与系统科学 2021-06-01 Vipin Chandra Pandey , Nikhil Gupta , K. R. Niazi , Anil Swarnkar , Rayees Ahmad Thokar

Understanding and predicting the electricity demand responses to prices are critical activities for system operators, retailers, and regulators. While conventional machine learning and time series analyses have been adequate for the routine…

信号处理 · 电气工程与系统科学 2024-10-07 Adrian Esteban-Perez , Derek Bunn , Yashar Ghiassi-Farrokhfal

Existing AI alignment approaches assume that preferences are static, which is unrealistic: our preferences change, and may even be influenced by our interactions with AI systems themselves. To clarify the consequences of incorrectly…

人工智能 · 计算机科学 2024-05-29 Micah Carroll , Davis Foote , Anand Siththaranjan , Stuart Russell , Anca Dragan

Nowadays, data-centers are largely under-utilized because resource allocation is based on reservation mechanisms which ignore actual resource utilization. Indeed, it is common to reserve resources for peak demand, which may occur only for a…

分布式、并行与集群计算 · 计算机科学 2018-07-03 Francesco Pace , Dimitrios Milios , Damiano Carra , Daniele Venzano , Pietro Michiardi

A Markov decision process can be parameterized by a transition kernel and a reward function. Both play essential roles in the study of reinforcement learning as evidenced by their presence in the Bellman equations. In our inquiry of various…

机器学习 · 计算机科学 2023-09-04 Falcon Z. Dai

This paper addresses a multi-echelon inventory management problem with a complex network topology where deriving optimal ordering decisions is difficult. Deep reinforcement learning (DRL) has recently shown potential in solving such…

机器学习 · 计算机科学 2024-01-30 Liqiang Cheng , Jun Luo , Weiwei Fan , Yidong Zhang , Yuan Li

This paper develops a practical framework for using observational data to audit the consumer surplus effects of AI-driven decisions, specifically in targeted pricing and algorithmic lending. Traditional approaches first estimate demand…

机器学习 · 统计学 2026-01-06 Zeyu Bian , Max Biggs , Ruijiang Gao , Zhengling Qi

Given data on the choices made by consumers for different offer sets, a key challenge is to develop parsimonious models that describe and predict consumer choice behavior while being amenable to prescriptive tasks such as pricing and…

机器学习 · 统计学 2025-04-15 Yanqiu Ruan , Xiaobo Li , Karthyek Murthy , Karthik Natarajan

Reinforcement learning (RL) agents have traditionally been tasked with maximizing the value function of a Markov decision process (MDP), either in continuous settings, with fixed discount factor $\gamma < 1$, or in episodic settings, with…

机器学习 · 计算机科学 2019-02-11 Silviu Pitis