Parrot:面向文本到图像生成的 Pareto 最优多奖励强化学习框架
计算机视觉与模式识别
2024-07-16 v2
摘要
近期研究表明,使用带有多个质量奖励的强化学习(RL)可以提升文本到图像(T2I)生成中图像的质量。然而,手动调整奖励权重存在挑战,并可能导致某些指标的过度优化。为解决此问题,我们提出 Parrot,通过多目标优化来处理该问题,并引入一种有效的多奖励优化策略以逼近 Pareto 最优。利用逐批 Pareto 最优选择,Parrot 自动识别不同奖励之间的最优权衡。我们使用新颖的多奖励优化算法联合优化 T2I 模型和提示扩展网络,从而显著提升图像质量,并允许在推理时通过奖励相关提示控制不同奖励的权衡。此外,我们在推理时引入以原始提示为中心的引导,确保提示扩展后对用户输入的忠实度。大量实验和用户研究验证了 Parrot 在多种质量标准(包括美学、人类偏好、图文对齐和图像情感)上优于多个基线方法。
引用
@article{arxiv.2401.05675,
title = {Parrot: Pareto-optimal Multi-Reward Reinforcement Learning Framework for Text-to-Image Generation},
author = {Seung Hyun Lee and Yinxiao Li and Junjie Ke and Innfarn Yoo and Han Zhang and Jiahui Yu and Qifei Wang and Fei Deng and Glenn Entis and Junfeng He and Gang Li and Sangpil Kim and Irfan Essa and Feng Yang},
journal= {arXiv preprint arXiv:2401.05675},
year = {2024}
}