保证最优策略存在的偏好关系条件
机器学习
2024-03-29 v2
摘要
从偏好反馈中学习(LfPF)在训练大语言模型(LLM)以及某些类型的交互式学习智能体中发挥着重要作用。然而,LfPF 算法的理论与应用之间存在显著差距。当前保证 LfPF 问题中最优策略存在的结论假设偏好与转移动态均由马尔可夫决策过程决定。我们引入了直接偏好过程(Direct Preference Process),一种用于分析部分可观测、非马尔可夫环境中 LfPF 问题的新框架。在该框架内,我们通过考察偏好的序结构,建立了保证最优策略存在的条件。我们表明,即使不存在能够表达学习目标的奖励函数,决策问题仍可具有由递归最优方程刻画的最优策略。这些发现凸显了探索不假设偏好由奖励生成的基于偏好的学习策略的必要性。
引用
@article{arxiv.2311.01990,
title = {Conditions on Preference Relations that Guarantee the Existence of Optimal Policies},
author = {Jonathan Colaço Carr and Prakash Panangaden and Doina Precup},
journal= {arXiv preprint arXiv:2311.01990},
year = {2024}
}
备注
v2: replaced with accepted AISTATS 2024 version, containing a new summary figure and one extra example. Results and conclusions are unchanged