Landmark-Assisted Monte Carlo Planning
Abstract
Landmarksconditions that must be satisfied at some point in every solution planhave contributed to major advancements in classical planning, but they have seldom been used in stochastic domains. We formalize probabilistic landmarks and adapt the UCT algorithm to leverage them as subgoals to decompose MDPs; core to the adaptation is balancing between greedy landmark achievement and final goal achievement. Our results in benchmark domains show that well-chosen landmarks can significantly improve the performance of UCT in online probabilistic planning, while the best balance of greedy versus long-term goal achievement is problem-dependent. The results suggest that landmarks can provide helpful guidance for anytime algorithms solving MDPs.
Cite
@article{arxiv.2508.11493,
title = {Landmark-Assisted Monte Carlo Planning},
author = {David H. Chan and Mark Roberts and Dana S. Nau},
journal= {arXiv preprint arXiv:2508.11493},
year = {2025}
}
Comments
To be published in the Proceedings of the 28th European Conference on Artificial Intelligence