RoDiF:面向受损人类反馈的鲁棒直接微调扩散策略
机器人学
2026-02-03 v1 机器学习
摘要
扩散策略是机器人控制的强大范式,但通过人类偏好进行微调面临根本性挑战,由于去噪过程的多步骤结构。为克服这一问题,我们引入统一的马尔可夫决策过程(MDP)表述,将扩散去噪链与环境动力学一致集成,实现扩散策略的无奖励直接偏好优化(DPO)。在此表述基础上,我们提出RoDiF(鲁棒直接微调),该方法显式处理受损的人类偏好。RoDiF通过几何假设切割视角重新解释DPO目标,并采用保守切割策略以实现鲁棒性,而无需假设特定的噪声分布。大量实验在长期操作任务上表明,RoDiF consistently outperforms state-of-the-art baselines, effectively steering pretrained diffusion policies of diverse architectures to human-preferred modes, while maintaining strong performance even under 30% corrupted preference labels.
引用
@article{arxiv.2602.00886,
title = {RoDiF: Robust Direct Fine-Tuning of Diffusion Policies with Corrupted Human Feedback},
author = {Amitesh Vatsa and Zhixian Xie and Wanxin Jin},
journal= {arXiv preprint arXiv:2602.00886},
year = {2026}
}