用于机组组合问题的强化学习
人工智能
2016-11-17 v1
摘要
本文通过将日前机组组合(UC)问题建模为马尔可夫决策过程(MDP),并寻找发电调度的低成本策略来求解该问题。我们提出了两种强化学习算法,并设计了第三种算法。我们将结果与先前使用模拟退火(SA)的工作进行了比较,结果显示运行成本降低了27%,运行时间为2.5分钟(而现有最先进方法需要2.5小时)。
引用
@article{arxiv.1507.05268,
title = {Reinforcement Learning for the Unit Commitment Problem},
author = {Gal Dalal and Shie Mannor},
journal= {arXiv preprint arXiv:1507.05268},
year = {2016}
}
备注
Accepted and presented in IEEE PES PowerTech, Eindhoven 2015, paper ID 462731