English

From Action Labels to Sets: Rethinking Action Supervision for Imitation Learning from Corrective Feedback

Robotics 2026-05-01 v3

Abstract

Behavior cloning (BC) optimizes policies by treating human demonstrations as pointwise action labels. While effective with accurate action labels, this formulation is brittle in practice: when human-provided actions are imperfect, treating each label as an exact target can steer the policy away from the underlying desired behavior, particularly when expressive models are used (e.g., energy-based models). As a result, we propose a human-in-the-loop alternative that replaces pointwise supervision with set-valued action targets. We introduce Contrastive policy Learning from Interactive Corrections (CLIC). CLIC leverages human corrections to construct and refine sets of desired actions, and optimizes a policy to place probability mass over these sets rather than over a single action target. This formulation naturally accommodates both absolute and relative corrections and can represent complex multi-modal behaviors. Extensive simulation and real-robot experiments show that the proposed approach leads to effective policy learning across diverse settings: CLIC remains competitive with the state of the art under accurate data while being substantially more robust under noisy, relative, and partial feedback. Our implementation is publicly available at https://clic-webpage.github.io/.

Keywords

Cite

@article{arxiv.2502.07645,
  title  = {From Action Labels to Sets: Rethinking Action Supervision for Imitation Learning from Corrective Feedback},
  author = {Zhaoting Li and Rodrigo Pérez-Dattari and Robert Babuska and Cosimo Della Santina and Jens Kober},
  journal= {arXiv preprint arXiv:2502.07645},
  year   = {2026}
}
R2 v1 2026-06-28T21:40:24.263Z